用于去除脚本标记的Python正则表达式

hcv*_*vst 3 python regex

我有点害怕问这个因为害怕报复"你无法用正则表达式解析HTML"邪教.为什么不re.subn(r'<(script).*?</\1>', '', data, re.DOTALL)删除多行"脚本",但最后只删除两个单行"脚本"?

谢谢,HC

>>> import re
>>> data = """\
<nothtml> 
  <head> 
    <title>Regular Expression HOWTO &mdash; Python v2.7.1 documentation</title> 
    <script type="text/javascript"> 
      var DOCUMENTATION_OPTIONS = {
        URL_ROOT:    '../',
        VERSION:     '2.7.1',
        COLLAPSE_MODINDEX: false,
        FILE_SUFFIX: '.html',
        HAS_SOURCE:  true
      };
    </script> 
    <script type="text/javascript" src="../_static/jquery.js"></script> 
    <script type="text/javascript" src="../_static/doctools.js"></script>
"""

>>> print (re.subn(r'<(script).*?</\1>', '', data, re.DOTALL)[0])
<nothtml> 
  <head> 
    <title>Regular Expression HOWTO &mdash; Python v2.7.1 documentation</title> 
    <script type="text/javascript"> 
      var DOCUMENTATION_OPTIONS = {
        URL_ROOT:    '../',
        VERSION:     '2.7.1',
        COLLAPSE_MODINDEX: false,
        FILE_SUFFIX: '.html',
        HAS_SOURCE:  true
      };
    </script> 
Run Code Online (Sandbox Code Playgroud)

Mar*_*air 6

撇开一般来说这是否是一个好主意的问题,你的例子的问题是第四个参数re.subn是count - flags在Python 2.6中没有参数,尽管它是作为Python 2.7中的第五个参数引入的.相反,你可以在正则表达式的末尾添加`(?s)以获得相同的效果:

>>> print (re.subn(r'<(script).*?</\1>(?s)', '', data)[0])

<nothtml> 
  <head> 
    <title>Regular Expression HOWTO &mdash; Python v2.7.1 documentation</title> 




>>>
Run Code Online (Sandbox Code Playgroud)

...或者如果您使用的是Python 2.7,这应该可行:

>>> print (re.subn(r'<(script).*?</\1>(?s)', '', 0, data)[0])
Run Code Online (Sandbox Code Playgroud)

...即0作为count参数插入.