我正在尝试从示例 pdf 中提取表格,问题是表格中没有用于分隔列的行。
这是文档的图片:
当我尝试运行时:
tables = camelot.read_pdf(filename)
print(tables[0].df)
Run Code Online (Sandbox Code Playgroud)
它打印正确的表格内容,但不考虑列,它将整个表格视为一列,如下所示:
0 Credit \nDate \nReference No. \nDescription...
1 09/10/2016 \n3194949206 \nOnline Banking tra...
2 09/20/2016 \n3194749206 \nOnline Banking tra...
3 09/20/2016 \n34236757678 \nCA TLR cash withd...
4 10/08/2016 \n5444 \nInterest \n ...
Run Code Online (Sandbox Code Playgroud)
当我运行这个时:
print(tables[0].df.shape)
Run Code Online (Sandbox Code Playgroud)
结果是 (5, 1)。
我尝试了另一种解决方案,指定如下所示的流:
tables = camelot.read_pdf(filename, flavor='stream')
Run Code Online (Sandbox Code Playgroud)
但随后它会获取并打印错误的数据,这就是结果:
0 1
0 Bank of Domino
1 Customer service information
2 P.O. Box 15001 1.888.DOMINO
3 Arlington, VA 18505 TDD/TTY users only: 1.800.288.4101
4 En Espanol: 1.800.688.6229
5 Account Number: 00000970987652 bankofdomino.com
6 ROBERT BELL Bank of Domino, N.A.
7 HOLLOW WAY,APT 503 P.O. Box 25125
8 SAN MESA, CA 92627-5125 San Mesa,CA 33390
Run Code Online (Sandbox Code Playgroud)
df 的形状为 (9, 2)。
我也尝试指定列 x 坐标但无济于事:
tables = camelot.read_pdf(filename, flavor='stream', columns="10, 120, 230, 470, 520, 650, 720, 800")
Run Code Online (Sandbox Code Playgroud)
它仍然得到错误的数据。
任何帮助表示赞赏。
提前致谢。
编辑:这是 pdf 样本。
https://mega.nz/file/cNtwDSpI#KuhG03P1Qg5kLa69jZ7ohb3FF8G6ITNuMFsFyQVMudw
好的,所以您有几个选择,我将用您的示例 pdf 文件为您提供一些示例:
tabula-py:import tabula
pdf_path = dir_path + "/bankk.pdf"
tb = tabula.read_pdf(pdf_path, pages='all')
Run Code Online (Sandbox Code Playgroud)
这将为您提供在 pdf 中检测到的所有表格的数据框列表,使用tb[0]将给出以下结果:
+----------+-------------+--------------------+---------+-------+----------+
| Date|Reference No.| Description| Credit| Debit| Balance|
+----------+-------------+--------------------+---------+-------+----------+
|09/10/2016| 3194949206|Online Banking tr...|$2,500.00| null|$14,000.49|
|09/20/2016| 3194749206|Online Banking tr...|$3,800.00| null|$17,800.49|
|09/20/2016| 34236757678|CA TLR cash withd...| null|$300.00|$17,500.49|
|10/08/2016| 5444| Interest| $0.64| null|$17,501.13|
+----------+-------------+--------------------+---------+-------+----------+
Run Code Online (Sandbox Code Playgroud)
pandas和pdfplumber:import pdfplumber
import pandas as pd
pdf = pdfplumber.open(pdf_path)
page = pdf.pages[0]
tb = page.extract_table(table_settings={"horizontal_strategy": "lines",
"vertical_strategy": "text",
"keep_blank_chars": "text",
"snap_tolerance": 5,})
df = pd.DataFrame(tb[1:], columns=tb[0])
Run Code Online (Sandbox Code Playgroud)
请注意,您可能需要使用
table_settings,要阅读更多信息,请参阅此链接
您可以使用该amazon-textract-textractor 包调用 texttract 并解析其输出。Textract 的优点是它适用于本机 pdf、扫描 pdf 或图像。例如您的图像:
from textractor import Textractor
from textractor.data.constants import TextractFeatures
extractor = Textractor(profile_name="default")
document = extractor.analyze_document(
file_source="./N6klE.png",
features=[TextractFeatures.TABLES],
)
Run Code Online (Sandbox Code Playgroud)
Textract 检测到两个表:
document.tables[1].to_pandas(use_columns=True)
Run Code Online (Sandbox Code Playgroud)
使用深度学习算法训练一个模型来检测pdf每一页的表格,然后用来获取表格数据,检测后获取表格数据pytesseract可以参考我在Medium上的文章: Image Table to DataFrame using Python光学字符识别
当然,您仍然可以使用camelot,但我更喜欢tabula-py简单的解决方案和深度学习更复杂的解决方案