ger*_*erm 2 html python pandas
我正在尝试使用pandas read_html函数阅读此处的 “众议院正式名单” 。
使用
df_list = pd.read_html('http://clerk.house.gov/member_info/olmbr.aspx',header=0,encoding = "UTF-8")
house = df_list[0]
Run Code Online (Sandbox Code Playgroud)
我确实得到了一个不错的DataFrame,其中包含代表姓名,州和地区。标头正确,编码也正确。到目前为止,一切都很好。
但是,问题在于聚会。没有派对的专栏。而是用字体(罗马或斜体)表示聚会。查看HTML源代码,这是民主人士的条目:
<tr><td><em>Adams, Alma S.</em></td><td>NC</td><td>12th</td></tr>
Run Code Online (Sandbox Code Playgroud)
这是共和党人的条目:
<tr><td>Anderholt, Robert B.</td><td>AL</td><td>4th</td></tr>
Run Code Online (Sandbox Code Playgroud)
共和党人<em></em>在他们的名字周围缺少标签。
人们将如何检索这一信息?可以用熊猫吗?还是需要一些更复杂的HTML解析器?如果是这样,哪个?
我认为您需要创建解析器:
import requests
from bs4 import BeautifulSoup
url = "http://clerk.house.gov/member_info/olmbr.aspx"
res = requests.get(url)
soup = BeautifulSoup(res.text,'html5lib')
table = soup.find_all('table')[0]
#print (table)
data = []
#remove first header
rows = table.find_all('tr')[1:]
for row in rows:
cols = row.find_all('td')
#get all children tags of first td
childrens = cols[0].findChildren()
#extracet all tags joined by ,
a = ', '.join([x.name for x in childrens]) if len(childrens) > 0 else ''
cols = [ele.text.strip() for ele in cols]
#add tag value for each row
cols.append(a)
data.append(cols)
Run Code Online (Sandbox Code Playgroud)
#DataFrame contructor
cols = ['Representative', 'State', 'District', 'Tag']
df = pd.DataFrame(data, columns=cols)
print (df.head())
Representative State District Tag
0 Abraham, Ralph Lee LA 5th
1 Adams, Alma S. NC 12th em
2 Aderholt, Robert B. AL 4th
3 Aguilar, Pete CA 31st em
4 Allen, Rick W. GA 12th
Run Code Online (Sandbox Code Playgroud)
也可以使用1和0为所有可能的标签创建列:
import requests
from bs4 import BeautifulSoup
url = "http://clerk.house.gov/member_info/olmbr.aspx"
res = requests.get(url)
soup = BeautifulSoup(res.text,'html5lib')
table = soup.find_all('table')[0]
#print (table)
data = []
rows = table.find_all('tr')[1:]
for row in rows:
cols = row.find_all('td')
childrens = cols[0].findChildren()
a = '|'.join([x.name for x in childrens]) if len(childrens) > 0 else ''
cols = [ele.text.strip() for ele in cols]
cols.append(a)
data.append(cols)
cols = ['Representative', 'State', 'District', 'Tag']
df = pd.DataFrame(data, columns=cols)
df = df.join(df.pop('Tag').str.get_dummies())
print (df.head())
Representative State District em strong
0 Abraham, Ralph Lee LA 5th 0 0
1 Adams, Alma S. NC 12th 1 0
2 Aderholt, Robert B. AL 4th 0 0
3 Aguilar, Pete CA 31st 1 0
4 Allen, Rick W. GA 12th 0 0
Run Code Online (Sandbox Code Playgroud)