我们可以在BeautifulSoup中使用xpath吗?

Shi*_*dla 93 python xpath urllib beautifulsoup

我正在使用BeautifulSoup来抓取一个网址,我有以下代码

import urllib
import urllib2
from BeautifulSoup import BeautifulSoup

url =  "http://www.example.com/servlet/av/ResultTemplate=AVResult.html"
req = urllib2.Request(url)
response = urllib2.urlopen(req)
the_page = response.read()
soup = BeautifulSoup(the_page)
soup.findAll('td',attrs={'class':'empformbody'})
Run Code Online (Sandbox Code Playgroud)

现在在上面的代码中我们可以findAll用来获取与它们相关的标签和信息,但我想使用xpath.是否可以将xpath与BeautifulSoup一起使用?如果可能的话,有人可以给我一个示例代码,以便更有帮助吗?

Mar*_*ers 147

Nope,BeautifulSoup本身不支持XPath表达式.

另一种库,LXML,不支持的XPath 1.0.它有一个BeautifulSoup兼容模式,它会像Soup一样尝试和解析破碎的HTML.但是,默认的lxml HTML解析器在解析损坏的HTML方面做得很好,我相信速度更快.

将文档解析为lxml树后,可以使用该.xpath()方法搜索元素.

try:
    # Python 2
    from urllib2 import urlopen
except ImportError:
    from urllib.request import urlopen
from lxml import etree

url =  "http://www.example.com/servlet/av/ResultTemplate=AVResult.html"
response = urlopen(url)
htmlparser = etree.HTMLParser()
tree = etree.parse(response, htmlparser)
tree.xpath(xpathselector)
Run Code Online (Sandbox Code Playgroud)

您可能感兴趣的是CSS Selector支持 ; 在lxml.html()类转换CSS语句转换为XPath表达式,使您的搜索response更加容易:

import lxml.html
import requests

url =  "http://www.example.com/servlet/av/ResultTemplate=AVResult.html"
response = requests.get(url, stream=True)
response.raw.decode_content = True
tree = lxml.html.parse(response.raw)
Run Code Online (Sandbox Code Playgroud)

完整的循环:BeautifulSoup本身确实有非常完整的CSS选择器支持:

from lxml.cssselect import CSSSelector

td_empformbody = CSSSelector('td.empformbody')
for elem in td_empformbody(tree):
    # Do something with these table cells.
Run Code Online (Sandbox Code Playgroud)

  • 证明消极是很难的; [BeautifulSoup 4文档](http://www.crummy.com/software/BeautifulSoup/bs4/doc/)具有搜索功能,"xpath"没有命中. (6认同)
  • 非常感谢 Pieters,我从你的代码中得到了两个信息,1。澄清我们不能将 xpath 与 BS 2.A 很好的例子一起使用 lxml。我们是否可以在特定文档中看到“我们无法以书面形式使用 BS 实现 xpath”,因为我们应该向那些要​​求澄清的人提供一些证据,对吗? (2认同)

Leo*_*son 101

我可以确认Beautiful Soup中没有XPath支持.

  • 注意:Leonard Richardson是Beautiful Soup的作者,因为你会看到你点击进入他的用户档案. (65认同)
  • 能够在BeautifulSoup中使用XPATH是非常好的 (18认同)
  • 那么替代品是什么? (2认同)
  • @leonard-richardson 现在是 2021 年了,您是否仍然确认 BeautifulSoup *仍然* 没有 xpath 支持? (2认同)

wor*_*ise 35

Martijn的代码不再正常工作(现在已经有4年多了......),该__CODE__行打印到控制台并且不会将值__CODE__赋给变量.参考这个,我能够使用请求和lxml找出这个工作:

from lxml import html
import requests

page = requests.get('http://econpy.pythonanywhere.com/ex/001.html')
tree = html.fromstring(page.content)
#This will create a list of buyers:
buyers = tree.xpath('//div[@title="buyer-name"]/text()')
#This will create a list of prices
prices = tree.xpath('//span[@class="item-price"]/text()')

print('Buyers: ', buyers)
print('Prices: ', prices)
Run Code Online (Sandbox Code Playgroud)


小智 15

BeautifulSoup有一个名为findNext的函数,来自当前元素导向的childern,所以:

father.findNext('div',{'class':'class_value'}).findNext('div',{'id':'id_value'}).findAll('a') 
Run Code Online (Sandbox Code Playgroud)

上面的代码可以模仿以下xpath:

div[class=class_value]/div[id=id_value]
Run Code Online (Sandbox Code Playgroud)


小智 14

from lxml import etree
from bs4 import BeautifulSoup
soup = BeautifulSoup(open('path of your localfile.html'),'html.parser')
dom = etree.HTML(str(soup))
print dom.xpath('//*[@id="BGINP01_S1"]/section/div/font/text()')
Run Code Online (Sandbox Code Playgroud)

以上使用了 Soup 对象与 lxml 的组合,可以使用 xpath 提取值