OCR REST API using Tesseract OCR Engine (via Tess4J).
3.7K
OCR REST API using Tesseract OCR Engine (via Tess4J)
Run with
docker run -it --rm -p 8000:8000 lgdd/tess4j-rest
You can navigate to /q/swagger-ui and test uploading an image.
Or you can quickly test the endpoint with curl:
curl -X 'POST' \
'http://localhost:8000/detect-text' \
-H 'accept: text/plain' \
-H 'Content-Type: multipart/form-data' \
-F '[email protected]'
# Parent folder path for tesseract data files
ENV TESSDATA_PREFIX="/opt/tesseract/tessdata"
# Suffix for the data repository to use.
# Either "best", "fast" or "".
# See https://github.com/tesseract-ocr/tessdata#readme
ENV TESSERACT_DATA_SUFFIX="best"
# Version of the data repository.
# See https://github.com/tesseract-ocr/tessdata#readme
ENV TESSERACT_DATA_VERSION="4.1.0"
# Additional languages to download on the application startup.
# For the possible values, see https://github.com/tesseract-ocr/tessdata
ENV TESSERACT_DATA_LANGS="fra,spa,deu"
For example, if you want to run this docker image and download additional Tesseract data for French, Spanish and German:
docker run -it --rm -p 8000:8000 -e TESSERACT_DATA_LANGS="fra,spa,deu" lgdd/tess4j-rest
And you should see the following logs:
INFO [com.git.lgd.tes.TesseractDataService] (vert.x-eventloop-thread-0) Downloading FRA data from https://github.com/tesseract-ocr/tessdata_best/raw/4.1.0/fra.traineddata
INFO [com.git.lgd.tes.TesseractDataService] (vert.x-eventloop-thread-0) Downloading SPA data from https://github.com/tesseract-ocr/tessdata_best/raw/4.1.0/spa.traineddata
INFO [com.git.lgd.tes.TesseractDataService] (vert.x-eventloop-thread-0) Downloading DEU data from https://github.com/tesseract-ocr/tessdata_best/raw/4.1.0/deu.traineddata
Readiness: /q/healh/ready
Liveness: /q/healh/live
Application is ready and live when all additional languages has been downloaded.
Content type
Image
Digest
sha256:f9471e682…
Size
271.7 MB
Last updated
almost 2 years ago
docker pull lgdd/tess4j-rest