Text extractor from documents as a service
201
Text extractor from documents as a service.
The service allows you to extract text from documents in PDF, DOC, DOCX and ODT formats.
All of the following examples will bring you the service on http://localhost:8080.
This will run the service with the default option:
docker run -d \
--name text-extractor \
-p 8080:80 \
sledgx/text-extractor
You can define log verbosity and maximum document size in the environment variables LOG_LEVEL and MAX respectively:
docker run -d \
--name text-extractor \
-e LOG_LEVEL=debug \
-e MAX_FILE_SIZE=15MB \
-p 8080:80 \
sledgx/text-extractor
The values accepted by LOG_LEVEL are error, warning, info, debug and notset, default is info.
The value of MAX_FILE_SIZE must be specified in human-readable format, i.e. the value followed by the unit of measurement (B, KB, MB, GB or TB).
The service exposes a single endpoint for text conversion.
With this method you can upload a document and receive the text contained in it as output. You can also specify whether to use the OCR system for extracting text from images, useful for documents in PDF format:
curl -X POST http://localhost:8080/convert \
-H 'Content-Type: multipart/form-data' \
-F '[email protected]' \
-F 'ocr_fallback=true'
or
curl -X POST http://localhost:8080/convert \
-H 'Content-Type: multipart/form-data' \
-H 'x-ocr-fallback: true' \
-F '[email protected]'
Released under the MIT License.
As with all Docker images, these likely also contain other software which may be under other licenses (such as Bash, etc from the base distribution, along with any direct or indirect dependencies of the primary software being contained).
Content type
Image
Digest
sha256:31158c747…
Size
532 MB
Last updated
over 3 years ago
docker pull sledgx/text-extractor