Shi*_*dla 93 python xpath urllib beautifulsoup
我正在使用BeautifulSoup来抓取一个网址,我有以下代码
import urllib
import urllib2
from BeautifulSoup import BeautifulSoup
url = "http://www.example.com/servlet/av/ResultTemplate=AVResult.html"
req = urllib2.Request(url)
response = urllib2.urlopen(req)
the_page = response.read()
soup = BeautifulSoup(the_page)
soup.findAll('td',attrs={'class':'empformbody'})
Run Code Online (Sandbox Code Playgroud)
现在在上面的代码中我们可以findAll用来获取与它们相关的标签和信息,但我想使用xpath.是否可以将xpath与BeautifulSoup一起使用?如果可能的话,有人可以给我一个示例代码,以便更有帮助吗?
Mar*_*ers 147
Nope,BeautifulSoup本身不支持XPath表达式.
另一种库,LXML,不支持的XPath 1.0.它有一个BeautifulSoup兼容模式,它会像Soup一样尝试和解析破碎的HTML.但是,默认的lxml HTML解析器在解析损坏的HTML方面做得很好,我相信速度更快.
将文档解析为lxml树后,可以使用该.xpath()方法搜索元素.
try:
# Python 2
from urllib2 import urlopen
except ImportError:
from urllib.request import urlopen
from lxml import etree
url = "http://www.example.com/servlet/av/ResultTemplate=AVResult.html"
response = urlopen(url)
htmlparser = etree.HTMLParser()
tree = etree.parse(response, htmlparser)
tree.xpath(xpathselector)
Run Code Online (Sandbox Code Playgroud)
您可能感兴趣的是CSS Selector支持 ; 在lxml.html()类转换CSS语句转换为XPath表达式,使您的搜索response更加容易:
import lxml.html
import requests
url = "http://www.example.com/servlet/av/ResultTemplate=AVResult.html"
response = requests.get(url, stream=True)
response.raw.decode_content = True
tree = lxml.html.parse(response.raw)
Run Code Online (Sandbox Code Playgroud)
完整的循环:BeautifulSoup本身确实有非常完整的CSS选择器支持:
from lxml.cssselect import CSSSelector
td_empformbody = CSSSelector('td.empformbody')
for elem in td_empformbody(tree):
# Do something with these table cells.
Run Code Online (Sandbox Code Playgroud)
Leo*_*son 101
我可以确认Beautiful Soup中没有XPath支持.
wor*_*ise 35
Martijn的代码不再正常工作(现在已经有4年多了......),该__CODE__行打印到控制台并且不会将值__CODE__赋给变量.参考这个,我能够使用请求和lxml找出这个工作:
from lxml import html
import requests
page = requests.get('http://econpy.pythonanywhere.com/ex/001.html')
tree = html.fromstring(page.content)
#This will create a list of buyers:
buyers = tree.xpath('//div[@title="buyer-name"]/text()')
#This will create a list of prices
prices = tree.xpath('//span[@class="item-price"]/text()')
print('Buyers: ', buyers)
print('Prices: ', prices)
Run Code Online (Sandbox Code Playgroud)
小智 15
BeautifulSoup有一个名为findNext的函数,来自当前元素导向的childern,所以:
father.findNext('div',{'class':'class_value'}).findNext('div',{'id':'id_value'}).findAll('a')
Run Code Online (Sandbox Code Playgroud)
上面的代码可以模仿以下xpath:
div[class=class_value]/div[id=id_value]
Run Code Online (Sandbox Code Playgroud)
小智 14
from lxml import etree
from bs4 import BeautifulSoup
soup = BeautifulSoup(open('path of your localfile.html'),'html.parser')
dom = etree.HTML(str(soup))
print dom.xpath('//*[@id="BGINP01_S1"]/section/div/font/text()')
Run Code Online (Sandbox Code Playgroud)
以上使用了 Soup 对象与 lxml 的组合,可以使用 xpath 提取值
| 归档时间: |
|
| 查看次数: |
110514 次 |
| 最近记录: |