A Docker image that automatically converts image-based PDFs to searchable PDFs.
476
This project creates a Docker image that automatically turns image-based (scanned) PDFs into searchable PDFs.
This is done primarily through the use of the excellent gkovacs/pdfocr ruby script.
In addition to that script, ghostscript is also used to force the page size to A4. Ideally that feature should be parameterised, but for now this is hard-coded.
This container requires you to create two volumes, and inbox and an outbox. Any PDF files dropped in the inbox are converted to searchable PDFs and dropped in the outbox.
Notes:
Example:
$ cd ~
$ mkdir ocr_in ocr_out
$ docker run -d --name pdfocr -v ~/ocr_in:/inbox -v ~/ocr_out:/outbox netservers/docker-pdfocr
Content type
Image
Digest
Size
257.4 MB
Last updated
over 10 years ago
docker pull netservers/docker-pdfocr