Sign inSign up

spikeyz/mindee-doctr-api

By spikeyz

•Updated 4 months ago

Mindee DocTR OCR Capabilities through REST api

Image
Machine learning & AI
0

408

spikeyz/mindee-doctr-api repository overview

⁠Github Repo

(https://github.com/spikeyz/mindee-doctr-api⁠)

⁠docTR REST API

A Docker image that wraps Mindee's docTR⁠ library and exposes all its features through an HTTP REST API built with FastAPI.

⁠Features

  • Full end-to-end OCR (image & PDF)
  • Text detection only
  • Text recognition only (from word crops)
  • Key Information Extraction (KIE)
  • Export as plain text, hOCR (XML), or searchable PDF
  • Annotated visualisation output
  • Runtime model selection across 9 detection and 8 recognition architectures
  • Model weights pre-downloaded at build time for fast cold starts
  • CPU and GPU (CUDA) support

⁠Endpoints

MethodPathDescription
GET/healthHealth check
GET/modelsList available model architectures
POST/ocrEnd-to-end OCR — image or PDF
POST/detectText detection only
POST/recognizeText recognition only (cropped word image)
POST/kieKey Information Extraction
POST/export/textOCR → plain text
POST/export/hocrOCR → hOCR XML
POST/export/searchable-pdfOCR → searchable PDF
POST/visualizeOCR → annotated PNG

Interactive documentation is available at http://localhost:8000/docs once the container is running.

⁠Quick start

⁠CPU (default)
docker compose build
docker compose up -d
⁠GPU (NVIDIA)

Requires nvidia-container-toolkit⁠.

Uncomment the args block in docker-compose.yml, then:

docker compose build --build-arg PIP_EXTRA_INDEX_URL=https://download.pytorch.org/whl/cu121
docker compose up -d

⁠Usage examples

# Health check
curl http://localhost:8000/health

# List available models
curl http://localhost:8000/models

# End-to-end OCR on an image
curl -X POST http://localhost:8000/ocr \
  -F "[email protected]"

# End-to-end OCR on a PDF with a different model pair
curl -X POST "http://localhost:8000/ocr?det_arch=fast_base&reco_arch=parseq" \
  -F "[email protected]"

# Detection only
curl -X POST http://localhost:8000/detect \
  -F "[email protected]"

# Recognition only (cropped word image)
curl -X POST http://localhost:8000/recognize \
  -F "file=@word_crop.png"

# Key Information Extraction
curl -X POST http://localhost:8000/kie \
  -F "[email protected]"

# Export as plain text
curl -X POST http://localhost:8000/export/text \
  -F "[email protected]"

# Export as hOCR XML
curl -X POST http://localhost:8000/export/hocr \
  -F "[email protected]" -o output.xml

# Export as searchable PDF
curl -X POST http://localhost:8000/export/searchable-pdf \
  -F "[email protected]" -o searchable.pdf

# Visualise with bounding boxes (returns PNG)
curl -X POST http://localhost:8000/visualize \
  -F "[email protected]" -o annotated.png

⁠Available models

⁠Detection
ArchitectureNotes
db_resnet34Lightweight DBNet
db_resnet50Default – best accuracy/speed trade-off
db_mobilenet_v3_largeMobile-friendly
linknet_resnet18Fast
linknet_resnet34
linknet_resnet50
fast_tinyFastest
fast_small
fast_base
⁠Recognition
ArchitectureNotes
crnn_vgg16_bnDefault
crnn_mobilenet_v3_smallLightweight
crnn_mobilenet_v3_large
sar_resnet31
master
vitstr_smallVision Transformer
vitstr_base
parseqState-of-the-art accuracy

⁠Configuration

Environment variableDefaultDescription
WORKERS1Number of uvicorn workers
DOCTR_CACHE_DIR/app/.cache/doctrModel weight cache directory

Tag summary

Content type

Image

Digest

sha256:b31c0e53e…

Size

3.1 GB

Last updated

4 months ago

docker pull spikeyz/mindee-doctr-api