如何使用camelot提取表格

azi*_*aon 2 python

我正在尝试从示例 pdf 中提取表格,问题是表格中没有用于分隔列的行。

这是文档的图片:

在此输入图像描述

当我尝试运行时:

tables = camelot.read_pdf(filename)
print(tables[0].df)
Run Code Online (Sandbox Code Playgroud)

它打印正确的表格内容,但不考虑列,它将整个表格视为一列,如下所示:

0  Credit  \nDate  \nReference No.  \nDescription...
1  09/10/2016  \n3194949206  \nOnline Banking tra...
2  09/20/2016  \n3194749206  \nOnline Banking tra...
3  09/20/2016  \n34236757678  \nCA TLR cash withd...
4  10/08/2016  \n5444  \nInterest  \n            ...
Run Code Online (Sandbox Code Playgroud)

当我运行这个时:

print(tables[0].df.shape)
Run Code Online (Sandbox Code Playgroud)

结果是 (5, 1)。

我尝试了另一种解决方案,指定如下所示的流:

tables = camelot.read_pdf(filename, flavor='stream')
Run Code Online (Sandbox Code Playgroud)

但随后它会获取并打印错误的数据,这就是结果:

                                0                                   1
0                  Bank of Domino                                    
1                                        Customer service information
2                  P.O. Box 15001                        1.888.DOMINO
3             Arlington, VA 18505  TDD/TTY users only: 1.800.288.4101
4                                          En Espanol: 1.800.688.6229
5  Account Number: 00000970987652                    bankofdomino.com
6                     ROBERT BELL                Bank of Domino, N.A.
7              HOLLOW WAY,APT 503                      P.O. Box 25125
8         SAN MESA, CA 92627-5125                   San Mesa,CA 33390
Run Code Online (Sandbox Code Playgroud)

df 的形状为 (9, 2)。

我也尝试指定列 x 坐标但无济于事:

tables = camelot.read_pdf(filename, flavor='stream', columns="10, 120, 230, 470, 520, 650, 720, 800")
Run Code Online (Sandbox Code Playgroud)

它仍然得到错误的数据。

任何帮助表示赞赏。

提前致谢。

编辑:这是 pdf 样本。

https://mega.nz/file/cNtwDSpI#KuhG03P1Qg5kLa69jZ7ohb3FF8G6ITNuMFsFyQVMudw

Lid*_*lef 5

好的,所以您有几个选择,我将用您的示例 pdf 文件为您提供一些示例:

选项 1,使用tabula-py:

import tabula

pdf_path = dir_path + "/bankk.pdf"
tb = tabula.read_pdf(pdf_path, pages='all')
Run Code Online (Sandbox Code Playgroud)

这将为您提供在 pdf 中检测到的所有表格的数据框列表,使用tb[0]将给出以下结果:

+----------+-------------+--------------------+---------+-------+----------+
|      Date|Reference No.|         Description|   Credit|  Debit|   Balance|
+----------+-------------+--------------------+---------+-------+----------+
|09/10/2016|   3194949206|Online Banking tr...|$2,500.00|   null|$14,000.49|
|09/20/2016|   3194749206|Online Banking tr...|$3,800.00|   null|$17,800.49|
|09/20/2016|  34236757678|CA TLR cash withd...|     null|$300.00|$17,500.49|
|10/08/2016|         5444|            Interest|    $0.64|   null|$17,501.13|
+----------+-------------+--------------------+---------+-------+----------+
Run Code Online (Sandbox Code Playgroud)

选项 2,使用pandas和pdfplumber:

import pdfplumber
import pandas as pd

pdf = pdfplumber.open(pdf_path)
page = pdf.pages[0]
tb = page.extract_table(table_settings={"horizontal_strategy": "lines",  
                                        "vertical_strategy": "text",
                                        "keep_blank_chars": "text",
                                         "snap_tolerance": 5,})
df = pd.DataFrame(tb[1:], columns=tb[0])
Run Code Online (Sandbox Code Playgroud)

请注意,您可能需要使用table_settings,要阅读更多信息,请参阅此链接

选项 3,使用 AWS Textract(Thomas编辑):

您可以使用该amazon-textract-textractor 包调用 texttract 并解析其输出。Textract 的优点是它适用于本机 pdf、扫描 pdf 或图像。例如您的图像:

from textractor import Textractor
from textractor.data.constants import TextractFeatures
extractor = Textractor(profile_name="default")
document = extractor.analyze_document(
    file_source="./N6klE.png",
    features=[TextractFeatures.TABLES],
)
Run Code Online (Sandbox Code Playgroud)

Textract 检测到两个表:

document.tables[1].to_pandas(use_columns=True)
Run Code Online (Sandbox Code Playgroud)

熊猫 df2

选项 4,使用深度学习

使用深度学习算法训练一个模型来检测pdf每一页的表格,然后用来获取表格数据,检测后获取表格数据pytesseract可以参考我在Medium上的文章: Image Table to DataFrame using Python光学字符识别


当然,您仍然可以使用camelot,但我更喜欢tabula-py简单的解决方案和深度学习更复杂的解决方案