Sign inSign up

datapsycho/oh-my-liteparser-cli-v2

By datapsycho

•Updated about 2 months ago

CLI version of only my liteparser to parse any document locally.

Image
Machine learning & AI
0

137

datapsycho/oh-my-liteparser-cli-v2 repository overview

⁠oh-my-liteparser-cli-v2

Document-intelligence CLI container image (Rust). Converts PDF/DOCX/DOC/PPTX/XLSX/image sources to Markdown via liteparse⁠, with selective OCR (only pages that actually need it). Unlike the companion oh-my-liteparser-lambda-v2 image, this one is meant to be pulled and run directly with docker run — no ECR, no AWS account, no cloud deploy step of any kind.

Base image: debian:bookworm-slim.

⁠Tags

{app_version}-liteparse{liteparse_version}, e.g. 2.3.0-liteparse2.11.1. latest is also published and kept up to date — since there's no deploy step pinning a digest the way a Lambda function does, latest here is a reasonable default for interactive use; pin an exact versioned tag instead for anything scripted/reproducible.


⁠Usage

docker pull datapsycho/oh-my-liteparser-cli-v2:latest

docker run --rm -v "$(pwd):/data" \
  datapsycho/oh-my-liteparser-cli-v2:latest \
  parse /data/report.docx -o /data/report.docx.md

Mount whatever host directory holds your input (and should receive the output) at /data, then reference paths inside the container as /data/.... The image's ENTRYPOINT is the cli binary itself, so no wrapper command is needed — arguments after the image name go straight to it.

⁠parse <FILE> [OPTIONS]
FlagDefaultMeaning
file (positional)—Path to the source document.
-o, --output <PATH>noneExact output file path. Wins outright over --output-dir/default naming.
--output-dir <PATH>noneDirectory to write the default-named file into. Ignored if --output is set.
--image-mode <off|placeholder|embed>placeholderHow embedded images are represented: dropped, ![](img_pN_M.png)-style reference, or inlined.
--no-linksoffEmit link text as plain text instead of [text](url).
--ocr <auto|always|never>autoauto: only OCR pages the pre-check flags. always/never: force.
--ocr-language <CODE>engTesseract language code, or a +-joined list (e.g. eng+fra) — see "OCR languages" below.
--plainoffSkip YAML front matter and page markers — raw concatenated Markdown only.
--password <PW>nonePassword for an encrypted PDF/Office source.

Default output filename (no --output): <original filename>.<ext>.md, e.g. report.docx → report.docx.md — preserves the source format in the name, written alongside the source or into --output-dir if given.

⁠is-complex <FILE>

Runs only the cheap text-layer pre-check (no OCR, no full parse) and prints the per-page verdict as JSON: needs_ocr, reasons (no-text, scanned, annotation-text, sparse-text, embedded-images, garbled, vector-text), text_length, text_coverage. Useful for routing/cost-estimation before committing to a full parse.

docker run --rm -v "$(pwd):/data" datapsycho/oh-my-liteparser-cli-v2:latest \
  is-complex /data/report.pdf

⁠OCR languages

The image ships Tesseract trained data for English plus the common Western European Latin-script languages: eng, fra, deu, spa, ita, por, nld (CJK is out of scope for this OCR engine). Set --ocr-language to a single code or your own +-joined subset (e.g. fra+deu) to force specific language(s) — mainly useful for a small speed gain when you already know the document's language; the CLI's own default is plain eng.

This limitation only applies when OCR actually runs. Whether a page needs OCR at all is decided purely by text density/images/cmap sanity, never by script or language — a Japanese, Arabic, or Russian source with a real embedded text layer skips OCR entirely and is extracted as-is. The limitation only bites a scanned/image-only non-Latin page, where Tesseract still produces output but confidently misreads non-Latin glyphs as something Latin — see low_confidence_pages below.

⁠Output: front matter and page markers

Every conversion (unless --plain) starts with a YAML front-matter block:

---
source: "report.docx"
filename: "report.docx"
pages: 5
ocr_enabled: true
ocr_engine: "tesseract"
ocr_pages_used: 3
low_confidence_pages: []
average_ocr_confidence: 0.870
detected_language: "fra"
detected_language_name: "French"
detected_script: "Latin"
detected_language_confidence: 1.000
detected_language_reliable: true
parsed_at: "2026-01-01T00:00:00.000Z"
---

ocr_pages_used counts pages OCR was attempted on, not pages whose text actually came from OCR — the engine discards any OCR result that overlaps text a page already has natively, so a page can count toward ocr_pages_used while its rendered Markdown is still 100% native text; check the per-page ocr_contributed_text marker (below) for the real answer.

average_ocr_confidence is a document-level mean of each OCR'd page's own confidence — a quick glance, not a substitute for low_confidence_pages: one badly-OCR'd page among many clean ones can still average out fine, so the per-page list stays the actual diagnostic. null when no page was OCR'd.

detected_language/_name/_script/_confidence/_reliable come from running whatlang⁠ over a bounded sample of the extracted text (first ~5 pages / ~3000 chars) — pure metadata, with no effect on --ocr-language/OCR routing, which stays entirely separate and under your control. detected_language/_name is which language (French, English, ...); detected_script is which writing system (Latin, Cyrillic, Arabic, ...) — many unrelated languages share a script, so detected_script alone can't distinguish English from French (both "Latin"), but it's a fast way to tell whether a document is even in a script this image's Latin-only tessdata can OCR at all. detected_language_reliable == false is still a real prediction, just low-confidence — surfaced, not hidden. All five detected_* fields are null together only when detection found nothing usable.

Each page is followed by a marker with a fixed key set (null, not omitted, when there's nothing to report):

<!-- page: {"page": N, "ocr": bool, "ocr_contributed_text": bool, "ocr_confidence": number | null} -->
ocrocr_contributed_textocr_confidenceMeaning
falsefalsenullNo OCR attempted — page didn't need it.
truefalsenullOCR attempted, but every result overlapped existing native text and was discarded — the page is still 100% native text.
truetrue0.61OCR attempted and its text survived into the page — the average Tesseract confidence across just that OCR-sourced text.

low_confidence_pages lists page numbers where ocr_contributed_text was true and the average confidence scored below 0.75 — a quality signal, not a language detector. Commonly caused by OCR running against the wrong --ocr-language (e.g. a non-Latin-script scan this image can't recognize at all), but a genuinely poor scan in the correct language can trip it too. Treat a flagged page as "verify before trusting," not a hard failure — the conversion still succeeds.

⁠Common errors

ErrorCause
password requiredSource is an encrypted PDF/Office document — pass --password.
OCR fails at parse time on an unrecognized language code--ocr-language doesn't match any .traineddata baked into this image (see "OCR languages" above).
LibreOffice conversion errors on .doc/.docx/etc.Source file is corrupt, or genuinely unsupported by LibreOffice's converter — not something this image's config can work around.

Tag summary

Content type

Image

Digest

sha256:f6c669fe3…

Size

248 MB

Last updated

about 2 months ago

docker pull datapsycho/oh-my-liteparser-cli-v2