使用 BeautifulSoup 将 HTML 插入到元素中

use*_*290 4 beautifulsoup python-3.x

当我尝试将以下 HTML 插入元素时

<div class="frontpageclass"><h3 id="feature_title">The Title</h3>... </div>
Run Code Online (Sandbox Code Playgroud)

bs4 正在像这样替换它:

<div class="frontpageclass">&lt;h3 id="feature_title"&gt;The Title &lt;/h3&gt;... &lt;div&gt;</div>
Run Code Online (Sandbox Code Playgroud)

我正在使用string,但它仍然弄乱了格式。

with open(html_frontpage) as fp:
   soup = BeautifulSoup(fp,"html.parser")

found_data = soup.find(class_= 'front-page__feature-image')
found_data.string = databasedata
Run Code Online (Sandbox Code Playgroud)

如果我尝试使用,found_data.string.replace_with我会收到 NoneType 错误。found_data是标签类型。

类似的问题,但他们使用的是 div,而不是类

Tom*_*lak 5

设置元素.textor.string导致值被 HTML 编码,这是正确的做法。它确保您插入的文本在浏览器中显示文档时以 1:1 的比例显示。

如果要插入实际的HTML,则需要在树中插入新节点。

from bs4 import BeautifulSoup

# always define a file encoding when working with text files
with open(html_frontpage, encoding='utf8') as fp:
    soup = BeautifulSoup(fp, "html.parser")

target = soup.find(class_= 'front-page__feature-image')

# empty out the target element if needed
target.clear()

# create a temporary document from your HTML
content = '<div class="frontpageclass"><h3 id="feature_title">The Title</h3>...</div>'
temp = BeautifulSoup(content)

# the nodes we want to insert are children of the <body> in `temp`
nodes_to_insert = temp.find('body').children

# insert them, in source order
for i, node in enumerate(nodes_to_insert):
    target.insert(i, node)
Run Code Online (Sandbox Code Playgroud)