CLI version of only my liteparser to parse any document locally.
137
Document-intelligence CLI container image (Rust). Converts PDF/DOCX/DOC/PPTX/XLSX/image
sources to Markdown via liteparse, with selective
OCR (only pages that actually need it). Unlike the companion oh-my-liteparser-lambda-v2
image, this one is meant to be pulled and run directly with docker run — no ECR, no AWS
account, no cloud deploy step of any kind.
Base image: debian:bookworm-slim.
{app_version}-liteparse{liteparse_version}, e.g. 2.3.0-liteparse2.11.1. latest is
also published and kept up to date — since there's no deploy step pinning a digest the
way a Lambda function does, latest here is a reasonable default for interactive use;
pin an exact versioned tag instead for anything scripted/reproducible.
docker pull datapsycho/oh-my-liteparser-cli-v2:latest
docker run --rm -v "$(pwd):/data" \
datapsycho/oh-my-liteparser-cli-v2:latest \
parse /data/report.docx -o /data/report.docx.md
Mount whatever host directory holds your input (and should receive the output) at
/data, then reference paths inside the container as /data/.... The image's
ENTRYPOINT is the cli binary itself, so no wrapper command is needed — arguments
after the image name go straight to it.
parse <FILE> [OPTIONS]| Flag | Default | Meaning |
|---|---|---|
file (positional) | — | Path to the source document. |
-o, --output <PATH> | none | Exact output file path. Wins outright over --output-dir/default naming. |
--output-dir <PATH> | none | Directory to write the default-named file into. Ignored if --output is set. |
--image-mode <off|placeholder|embed> | placeholder | How embedded images are represented: dropped, -style reference, or inlined. |
--no-links | off | Emit link text as plain text instead of [text](url). |
--ocr <auto|always|never> | auto | auto: only OCR pages the pre-check flags. always/never: force. |
--ocr-language <CODE> | eng | Tesseract language code, or a +-joined list (e.g. eng+fra) — see "OCR languages" below. |
--plain | off | Skip YAML front matter and page markers — raw concatenated Markdown only. |
--password <PW> | none | Password for an encrypted PDF/Office source. |
Default output filename (no --output): <original filename>.<ext>.md, e.g.
report.docx → report.docx.md — preserves the source format in the name, written
alongside the source or into --output-dir if given.
is-complex <FILE>Runs only the cheap text-layer pre-check (no OCR, no full parse) and prints the
per-page verdict as JSON: needs_ocr, reasons (no-text, scanned,
annotation-text, sparse-text, embedded-images, garbled, vector-text),
text_length, text_coverage. Useful for routing/cost-estimation before committing to
a full parse.
docker run --rm -v "$(pwd):/data" datapsycho/oh-my-liteparser-cli-v2:latest \
is-complex /data/report.pdf
The image ships Tesseract trained data for English plus the common Western European
Latin-script languages: eng, fra, deu, spa, ita, por, nld (CJK is out of
scope for this OCR engine). Set --ocr-language to a single code or your own
+-joined subset (e.g. fra+deu) to force specific language(s) — mainly useful for a
small speed gain when you already know the document's language; the CLI's own default
is plain eng.
This limitation only applies when OCR actually runs. Whether a page needs OCR at
all is decided purely by text density/images/cmap sanity, never by script or language —
a Japanese, Arabic, or Russian source with a real embedded text layer skips OCR
entirely and is extracted as-is. The limitation only bites a scanned/image-only
non-Latin page, where Tesseract still produces output but confidently misreads
non-Latin glyphs as something Latin — see low_confidence_pages below.
Every conversion (unless --plain) starts with a YAML front-matter block:
---
source: "report.docx"
filename: "report.docx"
pages: 5
ocr_enabled: true
ocr_engine: "tesseract"
ocr_pages_used: 3
low_confidence_pages: []
average_ocr_confidence: 0.870
detected_language: "fra"
detected_language_name: "French"
detected_script: "Latin"
detected_language_confidence: 1.000
detected_language_reliable: true
parsed_at: "2026-01-01T00:00:00.000Z"
---
ocr_pages_used counts pages OCR was attempted on, not pages whose text actually
came from OCR — the engine discards any OCR result that overlaps text a page already
has natively, so a page can count toward ocr_pages_used while its rendered Markdown
is still 100% native text; check the per-page ocr_contributed_text marker (below) for
the real answer.
average_ocr_confidence is a document-level mean of each OCR'd page's own
confidence — a quick glance, not a substitute for low_confidence_pages: one
badly-OCR'd page among many clean ones can still average out fine, so the per-page list
stays the actual diagnostic. null when no page was OCR'd.
detected_language/_name/_script/_confidence/_reliable come from running
whatlang over a bounded sample of the extracted
text (first ~5 pages / ~3000 chars) — pure metadata, with no effect on
--ocr-language/OCR routing, which stays entirely separate and under your control.
detected_language/_name is which language (French, English, ...); detected_script
is which writing system (Latin, Cyrillic, Arabic, ...) — many unrelated languages
share a script, so detected_script alone can't distinguish English from French (both
"Latin"), but it's a fast way to tell whether a document is even in a script this
image's Latin-only tessdata can OCR at all. detected_language_reliable == false is
still a real prediction, just low-confidence — surfaced, not hidden. All five
detected_* fields are null together only when detection found nothing usable.
Each page is followed by a marker with a fixed key set (null, not omitted, when
there's nothing to report):
<!-- page: {"page": N, "ocr": bool, "ocr_contributed_text": bool, "ocr_confidence": number | null} -->
ocr | ocr_contributed_text | ocr_confidence | Meaning |
|---|---|---|---|
false | false | null | No OCR attempted — page didn't need it. |
true | false | null | OCR attempted, but every result overlapped existing native text and was discarded — the page is still 100% native text. |
true | true | 0.61 | OCR attempted and its text survived into the page — the average Tesseract confidence across just that OCR-sourced text. |
low_confidence_pages lists page numbers where ocr_contributed_text was true and
the average confidence scored below 0.75 — a quality signal, not a language
detector. Commonly caused by OCR running against the wrong --ocr-language (e.g. a
non-Latin-script scan this image can't recognize at all), but a genuinely poor scan in
the correct language can trip it too. Treat a flagged page as "verify before trusting,"
not a hard failure — the conversion still succeeds.
| Error | Cause |
|---|---|
password required | Source is an encrypted PDF/Office document — pass --password. |
| OCR fails at parse time on an unrecognized language code | --ocr-language doesn't match any .traineddata baked into this image (see "OCR languages" above). |
LibreOffice conversion errors on .doc/.docx/etc. | Source file is corrupt, or genuinely unsupported by LibreOffice's converter — not something this image's config can work around. |
Content type
Image
Digest
sha256:f6c669fe3…
Size
248 MB
Last updated
about 2 months ago
docker pull datapsycho/oh-my-liteparser-cli-v2