Sign inSign up

datapsycho/oh-my-liteparser-cli

By datapsycho

•Updated 6 months ago

Developer-facing CLI image for parsing documents (PDF, DOCX, PPTX, XLSX, images) into txt and md.

Image
Machine learning & AI
Data science
0

93

datapsycho/oh-my-liteparser-cli repository overview

⁠oh-my-liteparser-cli — CLI Image

⁠Description

Developer-facing CLI image for parsing documents (PDF, DOCX, PPTX, XLSX, images) into multiple output formats. Built on top of the oh-my-liteparser base image — no extra setup required.

⁠Categories

document-parsing · cli · liteparse · pdf · ocr · markdown · nodejs


⁠Repository Overview

⁠What This Image Is

datapsycho/oh-my-liteparser-cli is the developer-facing command-line interface image for the oh-my-liteparser project. It extends the batteries-included base image (datapsycho/oh-my-liteparser) with the compiled CLI entrypoint, making it straightforward to parse documents from any supported format into rich page-aware Markdown, text-oriented markdown, LiteParse JSON, or screenshots — all from a single docker run command.

⁠What Problem It Solves

Running document parsing workflows locally often requires installing LibreOffice, ImageMagick, Tesseract, and Node.js with the right versions and environment variables. This image eliminates that setup entirely. Mount your input files and output directory, run the container, and get structured parsed output without touching the host environment.

⁠What Is Included

This image builds on datapsycho/oh-my-liteparser and adds:

  • The compiled oh-my-liteparser CLI (dist/cli/index.js)
  • Default entrypoint: node dist/cli/index.js
  • Default command: processes /data/sample → /output

All system-level dependencies (LibreOffice, ImageMagick, Ghostscript, Tesseract, LiteParse) are inherited from the base image.

⁠Supported Input Formats

PDF, DOCX, PPTX, XLSX, ODT, RTF, JPG, PNG, TIFF, WEBP, SVG, BMP, GIF — any format supported by LiteParse.

⁠Supported Output Formats
Flag valueDescriptionOutput suffix
plainText-oriented markdown with YAML frontmatter and page sections.txt.md
mdRich per-page markdown envelope with pages[] and merged content.md.json
liteparse-jsonFull LiteParse result wrapped with parse metadata.liteparse.json
screenshotPer-page PNG screenshots.screens/ directory
⁠How to Use It

Pull the image:

docker pull datapsycho/oh-my-liteparser-cli:1.5.2

Parse a single file:

docker run --rm \
  -v "$(pwd)/data:/data" \
  -v "$(pwd)/output:/output" \
  datapsycho/oh-my-liteparser-cli:1.5.2 \
  /data/sample/pdf/report.pdf --format md --output-dir /output

Parse an entire directory:

docker run --rm \
  -v "$(pwd)/data:/data" \
  -v "$(pwd)/output:/output" \
  datapsycho/oh-my-liteparser-cli:1.5.2 \
  /data/sample --format md --output-dir /output

Use the default command (processes /data/sample → /output):

docker run --rm \
  -v "$(pwd)/data:/data" \
  -v "$(pwd)/output:/output" \
  datapsycho/oh-my-liteparser-cli:1.5.2

CLI flags:

FlagDefaultDescription
<path>(required)File or directory to parse
--formatmdOutput format (plain, md, liteparse-json, screenshot)
--output-dir./outputDirectory to write results into
--ocroffEnable OCR for scanned documents
--confidenceunsetOCR confidence threshold between 0.0 and 1.0
⁠Tags
TagDescription
1.5.2Pinned release aligned with LiteParse v1.5.x
⁠Platform

linux/amd64 only.

⁠Credits

The document parsing capability in this image is powered by LiteParse⁠, developed by the LlamaIndex team. LiteParse provides structured, page-aware parsing across a wide range of document formats with optional OCR support. All credit for the underlying parsing engine goes to the LlamaIndex LiteParse contributors.

This CLI image packages LiteParse together with the oh-my-liteparser CLI layer and all required system dependencies, so developers can parse documents in a containerised environment without any local installation.

Tag summary

Content type

Image

Digest

sha256:d19c2f6cb…

Size

449.7 MB

Last updated

6 months ago

docker pull datapsycho/oh-my-liteparser-cli:1.5.2