如何在pytesseract中获得字符位置

Cha*_*lex 7 ocr image-processing python-2.7 python-tesseract pytesser

我试图使用pytesseract库获取图像文件的字符位置.

import pytesseract
from PIL import Image
print pytesseract.image_to_string(Image.open('5.png'))
Run Code Online (Sandbox Code Playgroud)

是否有任何图书馆可以获得角色的每个位置

小智 7

您尝试过使用 pytesseract.image_to_data() 吗?

data = pytesseract.image_to_data(img, output_type='dict')
boxes = len(data['level'])
for i in range(boxes ):
    (x, y, w, h) = (data['left'][i], data['top'][i], data['width'][i], data['height'][i])
    #Draw box        
    cv2.rectangle(img, (x, y), (x + w, y + h), (0, 255, 0), 2)
Run Code Online (Sandbox Code Playgroud)


小智 1

使用 pytesseract 似乎并不是获得该职位的最佳主意,但您可以这样做:

from pytesseract import pytesseract
pytesseract.run_tesseract('image.png', 'output', lang=None, boxes=False, config="hocr")
Run Code Online (Sandbox Code Playgroud)