par*_*cer 4 beautifulsoup html5lib web-scraping
我正在尝试安装html5lib.起初我试图安装最新版本(8或9个9),但它与我的BeautifulSoup发生冲突,所以我决定尝试更老的版本(0.9999999,7 个9).我安装了它,但是当我尝试使用它时:
>>> with urlopen("http://example.com/") as f:
document = html5lib.parse(f, encoding=f.info().get_content_charset())
Run Code Online (Sandbox Code Playgroud)
我收到一个错误:
Traceback (most recent call last):
File "<pyshell#11>", line 2, in <module>
document = html5lib.parse(f, encoding=f.info().get_content_charset())
File "C:\Python\Python35-32\lib\site-packages\html5lib\html5parser.py", line 35, in parse
return p.parse(doc, **kwargs)
File "C:\Python\Python35-32\lib\site-packages\html5lib\html5parser.py", line 235, in parse
self._parse(stream, False, None, *args, **kwargs)
File "C:\Python\Python35-32\lib\site-packages\html5lib\html5parser.py", line 85, in _parse
self.tokenizer = _tokenizer.HTMLTokenizer(stream, parser=self, **kwargs)
File "C:\Python\Python35-32\lib\site-packages\html5lib\_tokenizer.py", line 36, in __init__
self.stream = HTMLInputStream(stream, **kwargs)
File "C:\Python\Python35-32\lib\site-packages\html5lib\_inputstream.py", line 151, in HTMLInputStream
return HTMLBinaryInputStream(source, **kwargs)
TypeError: __init__() got an unexpected keyword argument 'encoding'
Run Code Online (Sandbox Code Playgroud)
有什么不对,我该怎么办?
我看到最新版本的html5lib中有关于bs4,html5lib.treebuilders._base已经不存在,usng bs4 4.4.1最新的兼容版本似乎是7个9的一个,一旦你安装它如下所示它工作正常:
pip3 install -U html5lib=="0.9999999"
Run Code Online (Sandbox Code Playgroud)
使用bs4 4.4.1测试:
In [1]: import bs4
In [2]: bs4.__version__
Out[2]: '4.4.1'
In [3]: import html5lib
In [4]: html5lib.__version__
Out[4]: '0.9999999'
In [5]: from urllib.request import urlopen
In [6]: with urlopen("http://example.com/") as f:
...: document = html5lib.parse(f, encoding=f.info().get_content_charset())
...:
In [7]:
Run Code Online (Sandbox Code Playgroud)
您可以在此提交中看到更改将treebuilders._base重命名为.base以反映名称已更改的公共状态:
您看到的错误是因为您仍在使用最新版本,在html5lib/_inputstream.py中,HTMLBinaryInputStream没有编码arg:
class HTMLBinaryInputStream(HTMLUnicodeInputStream):
"""Provides a unicode stream of characters to the HTMLTokenizer.
This class takes care of character encoding and removing or replacing
incorrect byte-sequences and also provides column and line tracking.
"""
def __init__(self, source, override_encoding=None, transport_encoding=None,
same_origin_parent_encoding=None, likely_encoding=None,
default_encoding="windows-1252", useChardet=True):
Run Code Online (Sandbox Code Playgroud)
设置override_encoding = f.info().get_content_charset()应该可以解决问题.
升级到最新版本的bs4也可以使用最新版本的html5lib:
In [16]: bs4.__version__
Out[16]: '4.5.1'
In [17]: html5lib.__version__
Out[17]: '0.999999999'
In [18]: with urlopen("http://example.com/") as f:
document = html5lib.parse(f, override_encoding=f.info().get_content_charset())
....:
In [19]:
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
4749 次 |
| 最近记录: |