USAGE: Extract text from documents or images of documents. Works best with PDF or TIFF, but can accept other formats. Outputs a txt file with the extracted text.
Will not work well on handwriting or images of docs (JPG, PNG, GIF)
Implementation of pytesseract (Python wrapper around Tesseract) using AWS S3 for input/output storage.
Content type
Image
Digest
Size
409 MB
Last updated
over 7 years ago
docker pull atiada/ocr-pytess