Developer-facing CLI image for parsing documents (PDF, DOCX, PPTX, XLSX, images) into txt and md.
93
Developer-facing CLI image for parsing documents (PDF, DOCX, PPTX, XLSX, images) into multiple output formats. Built on top of the oh-my-liteparser base image — no extra setup required.
document-parsing · cli · liteparse · pdf · ocr · markdown · nodejs
datapsycho/oh-my-liteparser-cli is the developer-facing command-line interface image for the oh-my-liteparser project. It extends the batteries-included base image (datapsycho/oh-my-liteparser) with the compiled CLI entrypoint, making it straightforward to parse documents from any supported format into rich page-aware Markdown, text-oriented markdown, LiteParse JSON, or screenshots — all from a single docker run command.
Running document parsing workflows locally often requires installing LibreOffice, ImageMagick, Tesseract, and Node.js with the right versions and environment variables. This image eliminates that setup entirely. Mount your input files and output directory, run the container, and get structured parsed output without touching the host environment.
This image builds on datapsycho/oh-my-liteparser and adds:
oh-my-liteparser CLI (dist/cli/index.js)node dist/cli/index.js/data/sample → /outputAll system-level dependencies (LibreOffice, ImageMagick, Ghostscript, Tesseract, LiteParse) are inherited from the base image.
PDF, DOCX, PPTX, XLSX, ODT, RTF, JPG, PNG, TIFF, WEBP, SVG, BMP, GIF — any format supported by LiteParse.
| Flag value | Description | Output suffix |
|---|---|---|
plain | Text-oriented markdown with YAML frontmatter and page sections | .txt.md |
md | Rich per-page markdown envelope with pages[] and merged content | .md.json |
liteparse-json | Full LiteParse result wrapped with parse metadata | .liteparse.json |
screenshot | Per-page PNG screenshots | .screens/ directory |
Pull the image:
docker pull datapsycho/oh-my-liteparser-cli:1.5.2
Parse a single file:
docker run --rm \
-v "$(pwd)/data:/data" \
-v "$(pwd)/output:/output" \
datapsycho/oh-my-liteparser-cli:1.5.2 \
/data/sample/pdf/report.pdf --format md --output-dir /output
Parse an entire directory:
docker run --rm \
-v "$(pwd)/data:/data" \
-v "$(pwd)/output:/output" \
datapsycho/oh-my-liteparser-cli:1.5.2 \
/data/sample --format md --output-dir /output
Use the default command (processes /data/sample → /output):
docker run --rm \
-v "$(pwd)/data:/data" \
-v "$(pwd)/output:/output" \
datapsycho/oh-my-liteparser-cli:1.5.2
CLI flags:
| Flag | Default | Description |
|---|---|---|
<path> | (required) | File or directory to parse |
--format | md | Output format (plain, md, liteparse-json, screenshot) |
--output-dir | ./output | Directory to write results into |
--ocr | off | Enable OCR for scanned documents |
--confidence | unset | OCR confidence threshold between 0.0 and 1.0 |
| Tag | Description |
|---|---|
1.5.2 | Pinned release aligned with LiteParse v1.5.x |
linux/amd64 only.
The document parsing capability in this image is powered by LiteParse, developed by the LlamaIndex team. LiteParse provides structured, page-aware parsing across a wide range of document formats with optional OCR support. All credit for the underlying parsing engine goes to the LlamaIndex LiteParse contributors.
This CLI image packages LiteParse together with the oh-my-liteparser CLI layer and all required system dependencies, so developers can parse documents in a containerised environment without any local installation.
Content type
Image
Digest
sha256:d19c2f6cb…
Size
449.7 MB
Last updated
6 months ago
docker pull datapsycho/oh-my-liteparser-cli:1.5.2