OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched.
OCRmyPDF requires ghostscript, tesseract an unpaper. We must connect these containers via the ssh protocol. The easiest solution is to use docker compose.
So create a file docker-compose.yml with content:
version: '3.4'
x-base: &base
volumes:
- .:/app
- ./tmp:/tmp
working_dir: /app
command: sshd
services:
ocrmypdf:
<<: *base
image: minidocks/ocrmypdf
links:
- tesseract
- unpaper
- gs
environment:
ALIAS_TESSERACT: ssh tesseract tesseract
ALIAS_UNPAPER: ssh unpaper unpaper
ALIAS_GS: ssh gs gs
gs:
<<: *base
image: minidocks/ghostscript
tesseract:
<<: *base
image: minidocks/tesseract:4-eng
environment:
OMP_THREAD_LIMIT: 1
unpaper:
<<: *base
image: minidocks/unpaper
And in the same directory run command:
docker-compose run --rm ocrmypdf -l eng input.pdf output.pdf
Content type
Image
Digest
Size
55.5 MB
Last updated
over 5 years ago
docker pull webuni/ocrmypdf