使用python和lxml模块从html中删除所有javascript标签和样式标签

joh*_*les 24 html python lxml

我正在使用http://lxml.de/库解析一个html文档.到目前为止,我已经想出如何从html文档中剥离标签在lxml中,如何删除标签但保留所有内容?但是该帖子中描述的方法会留下所有文本,剥离标签而不删除实际的脚本.我还发现了一类参考lxml.html.clean.Cleaner http://lxml.de/api/lxml.html.clean.Cleaner-class.html,但是这是明确的泥至于如何实际使用的类清理文件.任何帮助,也许是一个简短的例子对我有帮助!

acu*_*ich 57

下面是一个做你想做的事情的例子.对于HTML文档,这Cleaner是一个比使用更好的通用解决方案strip_elements,因为在这种情况下你想要剥离的不仅仅是<script>标签; 你也想摆脱onclick=function()其他标签上的属性.

#!/usr/bin/env python

import lxml
from lxml.html.clean import Cleaner

cleaner = Cleaner()
cleaner.javascript = True # This is True because we want to activate the javascript filter
cleaner.style = True      # This is True because we want to activate the styles & stylesheet filter

print("WITH JAVASCRIPT & STYLES")
print(lxml.html.tostring(lxml.html.parse('http://www.google.com')))
print("WITHOUT JAVASCRIPT & STYLES")
print(lxml.html.tostring(cleaner.clean_html(lxml.html.parse('http://www.google.com'))))
Run Code Online (Sandbox Code Playgroud)

您可以在lxml.html.clean.Cleaner文档中获取可以设置的选项列表; 你可以设置的一些选项True或False(默认)和其他选项如下:

cleaner.kill_tags = ['a', 'h1']
cleaner.remove_tags = ['p']
Run Code Online (Sandbox Code Playgroud)

注意kill vs remove之间的区别:

remove_tags:
  A list of tags to remove. Only the tags will be removed, their content will get pulled up into the parent tag.
kill_tags:
  A list of tags to kill. Killing also removes the tag's content, i.e. the whole subtree, not just the tag itself.
allow_tags:
  A list of tags to include (default include all).
Run Code Online (Sandbox Code Playgroud)

  • 我尝试了列表和元组,符号效果是相同的,标签不会被删除.经过一些进一步的研究后,我相信这是与ubuntu一起发布的lxml/html/clean.py版本中的一个错误.请注意http://lxml.de/api/lxml.html.clean-pysrc.html第253行的kill_tags在随附的clean.py版本中初始化为`kill_tags = set(self.kill_tags或())` Ubuntu刚刚初始化为`kill_tags = set()`.渲染它无效.谢谢,我会通知软件包维护者. (4认同)

Ash*_*her 6

以下是如何从 XML/HTML 树中删除和解析不同类型的 HTML 元素的一些示例。

关键建议:不依赖外部库并在“本机 python 2/3 代码”中完成所有操作是有帮助的。

以下是如何使用“本机”python 执行此操作的一些示例......

# (REMOVE <SCRIPT> to </script> and variations)
pattern = r'<[ ]*script.*?\/[ ]*script[ ]*>'  # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))

# (REMOVE HTML <STYLE> to </style> and variations)
pattern = r'<[ ]*style.*?\/[ ]*style[ ]*>'  # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))

# (REMOVE HTML <META> to </meta> and variations)
pattern = r'<[ ]*meta.*?>'  # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))

# (REMOVE HTML COMMENTS <!-- to --> and variations)
pattern = r'<[ ]*!--.*?--[ ]*>'  # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))

# (REMOVE HTML DOCTYPE <!DOCTYPE html to > and variations)
pattern = r'<[ ]*\![ ]*DOCTYPE.*?>'  # mach any char zero or more times
text = re.sub(pattern, '', text, flags=(re.IGNORECASE | re.MULTILINE | re.DOTALL))
Run Code Online (Sandbox Code Playgroud)

笔记:

re.IGNORECASE # is needed to match case sensitive <script> or <SCRIPT> or <Script>
re.MULTILINE # is needed to match newlines
re.DOTALL # is needed to match "special characters" and match "any character" 
Run Code Online (Sandbox Code Playgroud)

我已经在几个不同的 HTML 文件上对此进行了测试,包括 、 、 和 它工作“快速”并且可以跨换行!...

注意:它也不依赖于 beautifulsoup 或任何其他外部下载的库!

希望这可以帮助!

:)