Sign inSign up

nacrecontextlayer/nacre-parser

By nacrecontextlayer

•Updated 4 minutes ago

Nacre parser sidecar: bytes in, text and metadata out. PDF, Office, OpenDocument, EPUB, RTF, text.

Image
Machine learning & AI
Data science
Databases & storage
0

628

nacrecontextlayer/nacre-parser repository overview

⁠Nacre — parser sidecar

The Python sidecar a Nacre⁠ worker hands a document to: bytes in, {text, blocks, metadata} out. Text, Markdown, HTML and CSV; PDF through pdf-inspector; Word, PowerPoint, Excel, OpenDocument, EPUB and RTF through anydoc. A scanned PDF with no text layer is refused rather than indexed as nothing, and a declared type the bytes contradict is refused rather than sniffed around.

Part of the open core — github.com/nacre-work/nacre⁠, Apache 2.0. The index itself is nacrecontextlayer/nacre⁠; the worker reaches this process over HTTP and nothing else does.

⁠Configuration

variabledefaultmeaning
PORT8090where it listens
NACRE_PARSER_ALLOW_PRIVATE_URLSunsettrue lets a document ingested by URL point at a private address. Off by default, because a tenant must not be able to point this process at the cloud metadata endpoint

Two dependencies, both pinned, and the reason is written in the repository's services/parser/requirements.txt: this process runs hostile input, so its dependency surface is the whole decision. OCR is deliberately not enabled — the extractor's other mode sends the document to a hosted service, and the airgapped profile rests on this process reaching nothing.

Tags: {version} and latest, amd64 and arm64, always the same version as nacre. ghcr.io/nacre-work/nacre-parser is the canonical address; this repository is a mirror pushed by the same build at the same tags.

nacre.work⁠ · Quickstart⁠ · Configuration⁠ · Releases⁠

Tag summary

Content type

Image

Digest

sha256:b52d8950b…

Size

54.3 MB

Last updated

4 minutes ago

docker pull nacrecontextlayer/nacre-parser