使用 python 和 pdfkit 将 PDF 转换为 HTML

Duz*_*uzy 8 python pdf-to-html pdfkit

Adobe在此网站上写了有关使用 pdfkit 将 pdf 转换为 html 的内容

\n

他们使用pdfkit.from_pdf(...)方法。

\n
\n

此脚本使用 \xe2\x80\x98pdfkit\xe2\x80\x99 库将 PDF 文件转换为 HTML。要使用此脚本,您需要安装 \xe2\x80\x98pdfkit\xe2\x80\x99 库...

\n
\n

当我想使用这个方法时出现错误

\n
Traceback (most recent call last):\n  File "C:\\TestPdfToHtml\\script.py", line 7, in <module>\n    html_file = pdfkit.from_pdf(pdf_file, "my_html_file.html")\n                ^^^^^^^^^^^^^^^\nAttributeError: module \'pdfkit\' has no attribute \'from_pdf\'. Did you mean: \'from_url\'?\n
Run Code Online (Sandbox Code Playgroud)\n

我该如何解决这个问题?

\n

下面是完整的脚本

\n
import pdfkit\n# Read the PDF file\npdf_file = open(\'test2.pdf\', \'rb\')\n# Convert the PDF to HTML\nhtml_file = pdfkit.from_pdf(pdf_file, "my_html_file.html")\n# Close the PDF file\npdf_file.close()\n
Run Code Online (Sandbox Code Playgroud)\n

小智 -7

也许较新版本的 pdfkit 不支持 pdfkit.from_pdf。您可以尝试 pdfkit.from_file()

pdfkit.from_file(pdf_file, html_file)
Run Code Online (Sandbox Code Playgroud)

希望这可以帮助。

  • 投反对票。我很确定你没有尝试过这个。这不起作用。当我使用实际的 PDF 输入文件(即使是 pdfkit 本身生成的极其简单的文件)尝试此操作时,我得到:“wkhtmltopdf 以非零代码 1 退出。错误:由于未知错误,以代码 1 退出。”我'我很确定 pdfkit 不能做到这一点。我想知道为什么 Adob​​e 发布了这些废话...... (4认同)