如何通过BeautifulSoup提取后通过正则表达式运行属性值?

Joh*_*ohn 0 python regex unicode url beautifulsoup

我有一个URL,我想解析其中的一部分,特别是widgetid:

<a href="http://www.somesite.com/process.asp?widgetid=4530">Widgets Rock!</a>
Run Code Online (Sandbox Code Playgroud)

我写过这篇Python(我在Python上有点新手 - 版本是2.7):

import re
from bs4 import BeautifulSoup

doc = open('c:\Python27\some_xml_file.txt')
soup = BeautifulSoup(doc)


links = soup.findAll('a')

# debugging statements

print type(links[7])
# output: <class 'bs4.element.Tag'>

print links[7]
# output: <a href="http://www.somesite.com/process.asp?widgetid=4530">Widgets Rock!</a>

theURL = links[7].attrs['href']
print theURL
# output: http://www.somesite.com/process.asp?widgetid=4530

print type(theURL)
# output: <type 'unicode'>

is_widget_url = re.compile('[0-9]')
print is_widget_url.match(theURL)
# output: None (I know this isn't the correct regex but I'd think it
#         would match if there's any number in there!)
Run Code Online (Sandbox Code Playgroud)

我认为我缺少正则表达式(或我对如何使用它们的理解),但我无法弄明白.

谢谢你的帮助!

Dan*_*man 5

这个问题与BeautifulSoup没有任何关系.

问题是,正如文档所解释的那样,match只在字符串开头匹配.由于您要查找的数字位于字符串的末尾,因此不返回任何内容.

要在任何地方匹配数字,请使用search- 您可能希望将\d实体用于数字.

matches = re.search(r'\d+', theURL)
Run Code Online (Sandbox Code Playgroud)