我正在使用http://lxml.de/库解析一个html文档.到目前为止,我已经想出如何从html文档中剥离标签在lxml中,如何删除标签但保留所有内容?但是该帖子中描述的方法会留下所有文本,剥离标签而不删除实际的脚本.我还发现了一类参考lxml.html.clean.Cleaner http://lxml.de/api/lxml.html.clean.Cleaner-class.html,但是这是明确的泥至于如何实际使用的类清理文件.任何帮助,也许是一个简短的例子对我有帮助!
acu*_*ich 57
下面是一个做你想做的事情的例子.对于HTML文档,这Cleaner是一个比使用更好的通用解决方案strip_elements,因为在这种情况下你想要剥离的不仅仅是<script>标签; 你也想摆脱onclick=function()其他标签上的属性.
#!/usr/bin/env python
import lxml
from lxml.html.clean import Cleaner
cleaner = Cleaner()
cleaner.javascript = True # This is True because we want to activate the javascript filter
cleaner.style = True # This is True because we want to activate the styles & stylesheet filter
print("WITH JAVASCRIPT & STYLES")
print(lxml.html.tostring(lxml.html.parse('http://www.google.com')))
print("WITHOUT JAVASCRIPT & STYLES")
print(lxml.html.tostring(cleaner.clean_html(lxml.html.parse('http://www.google.com'))))
Run Code Online (Sandbox Code Playgroud)
您可以在lxml.html.clean.Cleaner文档中获取可以设置的选项列表; 你可以设置的一些选项True或False(默认)和其他选项如下:
cleaner.kill_tags = ['a', 'h1']
cleaner.remove_tags = ['p']
Run Code Online (Sandbox Code Playgroud)
注意kill vs remove之间的区别:
remove_tags:
A list of tags to remove. Only the tags will be removed, their content will get pulled up into the parent tag.
kill_tags:
A list of tags to kill. Killing also removes the tag's content, i.e. the whole subtree, not just the tag itself.
allow_tags:
A list of tags to include (default include all).
Run Code Online (Sandbox Code Playgroud)
以下是如何从 XML/HTML 树中删除和解析不同类型的 HTML 元素的一些示例。
关键建议:不依赖外部库并在“本机 python 2/3 代码”中完成所有操作是有帮助的。
以下是如何使用“本机”python 执行此操作的一些示例......
# (REMOVE <SCRIPT> to </script> and variations)
pattern = r'<[ ]*script.*?\/[ ]*script[ ]*>' # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))
# (REMOVE HTML <STYLE> to </style> and variations)
pattern = r'<[ ]*style.*?\/[ ]*style[ ]*>' # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))
# (REMOVE HTML <META> to </meta> and variations)
pattern = r'<[ ]*meta.*?>' # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))
# (REMOVE HTML COMMENTS <!-- to --> and variations)
pattern = r'<[ ]*!--.*?--[ ]*>' # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))
# (REMOVE HTML DOCTYPE <!DOCTYPE html to > and variations)
pattern = r'<[ ]*\![ ]*DOCTYPE.*?>' # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))
Run Code Online (Sandbox Code Playgroud)
笔记:
re.IGNORECASE # is needed to match case sensitive <script> or <SCRIPT> or <Script>
re.MULTILINE # is needed to match newlines
re.DOTALL # is needed to match "special characters" and match "any character"
Run Code Online (Sandbox Code Playgroud)
我已经在几个不同的 HTML 文件上对此进行了测试,包括 、 、 和 它工作“快速”并且可以跨换行!...
注意:它也不依赖于 beautifulsoup 或任何其他外部下载的库!
希望这可以帮助!
:)
| 归档时间: |
|
| 查看次数: |
16162 次 |
| 最近记录: |