Nacre parser sidecar: bytes in, text and metadata out. PDF, Office, OpenDocument, EPUB, RTF, text.
628
The Python sidecar a Nacre worker hands a document to: bytes in, {text, blocks, metadata} out. Text, Markdown, HTML and CSV; PDF through pdf-inspector; Word, PowerPoint, Excel, OpenDocument, EPUB and RTF through anydoc. A scanned PDF with no text layer is refused rather than indexed as nothing, and a declared type the bytes contradict is refused rather than sniffed around.
Part of the open core — github.com/nacre-work/nacre, Apache 2.0. The index itself is nacrecontextlayer/nacre; the worker reaches this process over HTTP and nothing else does.
| variable | default | meaning |
|---|---|---|
PORT | 8090 | where it listens |
NACRE_PARSER_ALLOW_PRIVATE_URLS | unset | true lets a document ingested by URL point at a private address. Off by default, because a tenant must not be able to point this process at the cloud metadata endpoint |
Two dependencies, both pinned, and the reason is written in the repository's services/parser/requirements.txt: this process runs hostile input, so its dependency surface is the whole decision. OCR is deliberately not enabled — the extractor's other mode sends the document to a hosted service, and the airgapped profile rests on this process reaching nothing.
Tags: {version} and latest, amd64 and arm64, always the same version as nacre. ghcr.io/nacre-work/nacre-parser is the canonical address; this repository is a mirror pushed by the same build at the same tags.
Content type
Image
Digest
sha256:b52d8950b…
Size
54.3 MB
Last updated
4 minutes ago
docker pull nacrecontextlayer/nacre-parser