Sign inSign up

datapsycho/oh-my-liteparser-lambda-v2

By datapsycho

•Updated about 2 months ago

Document-intelligence Lambda container image (Rust). via liteparse.

Image
Machine learning & AI
0

224

datapsycho/oh-my-liteparser-lambda-v2 repository overview

⁠oh-my-liteparser-lambda-v2

Document-intelligence Lambda container image (Rust). Converts PDF/DOCX/DOC/image sources to Markdown via liteparse⁠, with selective OCR (only pages that actually need it), and is invoked directly with an S3-to-S3 config — no API Gateway/HTTP involved.

Architecture: arm64 only. Base image: public.ecr.aws/lambda/provided:al2023.

⁠Requirements

  • Graviton (arm64) Lambda — not optional. This image is built and only published for arm64. The Lambda function's --architectures must be set to arm64; there is no x86_64 build. Deploying with --architectures x86_64 (or letting it default) fails every invoke with Runtime.InvalidEntrypoint/ProcessSpawnFailed — this is an architecture mismatch, not a bug to work around, and switching to x86_64 is not a supported alternative for this image.
  • An IAM execution role for the function (S3 access) and, separately, an invoke-only IAM policy for whoever calls it — see "DevOps guideline" and "Developer guideline" below.
  • Your own Amazon ECR repository — Lambda cannot run this image straight from Docker Hub (see next section).

⁠Tags

{app_version}-liteparse{liteparse_version}, e.g. 2.0.0-liteparse2.11.1 — always deploy against an exact tag like this, never latest. latest is published as a convenience pointer for docker pull only and always tracks the newest release; it is not a stable, reproducible target and should never be what a live Lambda function points at.


⁠DevOps guideline: deploying this image

⁠Important: this image cannot run directly from Docker Hub

AWS Lambda only pulls container images from Amazon ECR — not Docker Hub. To deploy this image, pull it here, then push it into your own ECR before pointing Lambda at it:

docker pull datapsycho/oh-my-liteparser-lambda-v2:<TAG>

AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
AWS_REGION=<your-region>
ECR_URI="${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/oh-my-liteparser-lambda-v2"

aws ecr create-repository --repository-name oh-my-liteparser-lambda-v2 --region "${AWS_REGION}"   # first time only

docker tag datapsycho/oh-my-liteparser-lambda-v2:<TAG> "${ECR_URI}:<TAG>"

aws ecr get-login-password --region "${AWS_REGION}" \
  | docker login --username AWS --password-stdin "${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com"

docker push "${ECR_URI}:<TAG>"
⁠Create the execution role (once, if you don't already have one)

Trusted for lambda.amazonaws.com, with exactly the S3 access this function needs — scope Resource to your actual bucket(s), not *:

{
  "Version": "2012-10-17",
  "Statement": [
    { "Sid": "ReadSourceDocuments", "Effect": "Allow", "Action": ["s3:GetObject", "s3:HeadObject"], "Resource": "arn:aws:s3:::<your-bucket>/*" },
    { "Sid": "WriteDestinationOutput", "Effect": "Allow", "Action": ["s3:PutObject"], "Resource": "arn:aws:s3:::<your-bucket>/*" },
    { "Sid": "ListBucket", "Effect": "Allow", "Action": ["s3:ListBucket", "s3:GetBucketLocation"], "Resource": "arn:aws:s3:::<your-bucket>" }
  ]
}

Plus the AWS-managed AWSLambdaBasicExecutionRole policy, for CloudWatch Logs.

⁠Create the function

--architectures arm64 must match this image's actual architecture — mismatching it fails every invoke with Runtime.InvalidEntrypoint/ProcessSpawnFailed.

aws lambda create-function \
  --function-name <your-function-name> \
  --package-type Image \
  --code ImageUri="${ECR_URI}:<TAG>" \
  --role arn:aws:iam::${AWS_ACCOUNT_ID}:role/<your-execution-role> \
  --timeout 300 \
  --memory-size 3008 \
  --architectures arm64 \
  --region "${AWS_REGION}"

300s timeout / 3008MB memory are verified comfortable for real documents up to 362 pages with mixed OCR (~107s duration, ~306MB peak memory). Raise either later via update-function-configuration if your documents are larger or more OCR-heavy — no image rebuild needed. /tmp (ephemeral storage) defaults to 512MB; raise it (--ephemeral-storage Size=<MB>) if you expect very large or image-heavy source documents.

VariableExamplePurpose
ALLOWED_SOURCE_BUCKETmy-bucketReject any invoke naming a different source bucket
ALLOWED_SOURCE_PREFIXinput/Reject any invoke whose source key isn't under this prefix
ALLOWED_DEST_BUCKETmy-bucketReject any invoke naming a different destination bucket
ALLOWED_DEST_PREFIXoutput/Reject any invoke whose destination isn't under this prefix
AWS_LAMBDA_LOG_FORMATJSONStructured JSON log lines (queryable in CloudWatch Logs Insights)
AWS_LAMBDA_LOG_LEVEL / RUST_LOGINFO / infoLog verbosity

These guardrails are deploy-time only — not part of the invoke payload, so a caller can't loosen them. A request whose destination folder would equal its own source folder is always rejected regardless of these settings (circular-write protection).

aws lambda update-function-configuration \
  --function-name <your-function-name> \
  --region "${AWS_REGION}" \
  --environment 'Variables={ALLOWED_SOURCE_BUCKET=<bucket>,ALLOWED_DEST_BUCKET=<bucket>,ALLOWED_SOURCE_PREFIX=input/,ALLOWED_DEST_PREFIX=output/,AWS_LAMBDA_LOG_FORMAT=JSON,RUST_LOG=info}' \
  --logging-config LogFormat=JSON,ApplicationLogLevel=INFO,SystemLogLevel=INFO
⁠Subsequent updates

Repoint at a new tag — nothing else changes:

aws lambda update-function-code \
  --function-name <your-function-name> \
  --image-uri "${ECR_URI}:<NEW_TAG>" \
  --region "${AWS_REGION}"

⁠Developer guideline: invoking the deployed function

⁠Invoke policy

The caller (a service, a scheduled job, a human via aws lambda invoke) needs only lambda:InvokeFunction on this one function's ARN — no S3 permissions at all; the execution role above handles all S3 access on the caller's behalf. Ask your DevOps/infra team to attach this to your IAM user/role:

{
  "Version": "2012-10-17",
  "Statement": [
    { "Sid": "InvokeDocumentConversion", "Effect": "Allow", "Action": "lambda:InvokeFunction", "Resource": "arn:aws:lambda:<region>:<account-id>:function:<your-function-name>" }
  ]
}

Scoped to that one function ARN, not function:*, so it can't be used to invoke anything else in the account.

⁠Invoke payload
{
  "source_bucket": "my-input-bucket",
  "source_key": "input/report.docx",
  "destination_bucket": "my-output-bucket",
  "destination_prefix": "output/",      // optional, default ""
  "destination_key": null,              // optional — overrides prefix + derived name
  "image_mode": "placeholder",          // optional — off | placeholder | embed
  "extract_links": true,                // optional, default true
  "ocr_mode": "auto",                   // optional — auto | always | never
  "ocr_language": null,                 // optional — see "OCR languages" below
  "password": null,                     // optional — encrypted PDF/Office sources
  "plain": false                        // optional — skip front matter/page markers
}

Only source_bucket/source_key/destination_bucket are required; everything else has a default. source_bucket/destination_bucket (and prefixes, if ALLOWED_* guardrails are set) must satisfy whatever the deploying account configured — ask DevOps which bucket/prefix you're allowed to use.

⁠OCR languages

The image ships Tesseract trained data for English plus the common Western European Latin-script languages: eng, fra, deu, spa, ita, por, nld (CJK is out of scope for this OCR engine). Leave ocr_language unset and all of these are recognized together in a single pass — Tesseract loads every dictionary at once and picks the best-matching one per word/region, so you don't need to know or specify a language up front. Set ocr_language to a single code (or your own +-joined subset, e.g. "fra+deu") to force specific language(s) — typically only useful for a small speed gain when you already know the document's language. An unrecognized code, or one without matching trained data in the image, fails the OCR step at invoke time.

This limitation only applies when OCR actually runs. Whether a page needs OCR at all is decided purely by text density/images/cmap sanity — never by script or language. A Japanese, Arabic, or Russian source with a real embedded text layer (not a scan) skips OCR entirely and its text is extracted as-is, so the Latin-only tessdata above is irrelevant to it. The limitation only bites a scanned/image-only non-Latin page, where OCR is the only way to get any text — Tesseract still produces output there, but confidently misreads non-Latin glyphs as something Latin, which is what the low-confidence signal below is for.

⁠Sample invoke
aws lambda invoke \
  --function-name <your-function-name> \
  --region <region> \
  --cli-binary-format raw-in-base64-out \
  --payload '{"source_bucket":"<bucket>","source_key":"input/report.docx","destination_bucket":"<bucket>","destination_prefix":"output/"}' \
  response.json && cat response.json
⁠Response
{
  "source": { "bucket": "...", "key": "..." },
  "destination": { "bucket": "...", "key": "..." },
  "pages": 5,
  "ocr_enabled": true,
  "ocr_engine": "tesseract",
  "ocr_pages_used": 3,
  "low_confidence_pages": [],
  "average_ocr_confidence": 0.87,
  "detected_language": "fra",
  "detected_language_name": "French",
  "detected_script": "Latin",
  "detected_language_confidence": 1.0,
  "detected_language_reliable": true
}

ocr_pages_used counts pages OCR was attempted on, not pages whose text actually came from OCR. The engine discards any OCR result that overlaps text a page already has natively, so a page can count toward ocr_pages_used while the Markdown that lands in S3 is still 100% native text for that page — check the per-page ocr_contributed_text marker (below) to know which pages' text genuinely came from OCR.

average_ocr_confidence is a document-level mean of each OCR'd page's own confidence — a quick glance, not a substitute for low_confidence_pages: a single badly-OCR'd page among many clean ones can still average out fine, so the per-page list stays the actual diagnostic. null when no page was OCR'd.

detected_language/_name/_script/_confidence/_reliable come from running whatlang⁠ over a bounded sample of the extracted text (first ~5 pages / ~3000 chars) — pure metadata with no effect on ocr_language/OCR routing, which stays entirely separate and under caller control. detected_language/_name is which language (French, English, ...); detected_script is which writing system (Latin, Cyrillic, Arabic, ...) — many unrelated languages share a script, so detected_script alone can't distinguish English from French (both "Latin"), but it's a fast way to tell whether a document is even in a script this image's Latin-only tessdata can OCR at all. detected_language_reliable == false is still a real prediction, just low-confidence — surfaced, not hidden. All five detected_* fields are null together only when detection found nothing usable.

The converted Markdown itself lands at destination.bucket/destination.key in S3 — this response is only the summary, not the content. Markdown includes a YAML front-matter block (source, filename, pages, ocr_enabled, ocr_engine, ocr_pages_used, low_confidence_pages, average_ocr_confidence, the five detected_* fields above, parsed_at) followed by per-page content with a marker — a JSON object with a fixed key set on every page (ocr_confidence is null, not omitted, when there's nothing to report), so it can be parsed with any JSON library rather than regex/string-split:

<!-- page: {"page": N, "ocr": bool, "ocr_contributed_text": bool, "ocr_confidence": number | null} -->
ocrocr_contributed_textocr_confidenceMeaning
falsefalsenullNo OCR attempted — page didn't need it.
truefalsenullOCR attempted, but every result overlapped existing native text and was discarded — the page is still 100% native text. ocr: true here does not mean the text you're reading came from OCR.
truetrue0.61OCR attempted and its text survived into the page — ocr_confidence is the average Tesseract confidence across just that OCR-sourced text.

(unless "plain": true was set, which omits front matter and markers entirely).

low_confidence_pages lists page numbers where ocr_contributed_text was true and the average confidence scored below the threshold (0.75) — a quality signal, not a language detector. It commonly catches OCR run against the wrong ocr_language (e.g. a non-Latin-script scan, which this image can't recognize at all — see "OCR languages" above), but a genuinely poor scan in the correct language can trip it too. Treat a flagged page as "don't trust this text blindly," not as a hard failure — the request still succeeds.

Caveat: the native-vs-OCR decision can differ across environments for the same input. Whether an OCR result is discarded (kept as native text) or kept (merged/overrides that region) comes down to a bounding-box overlap check inside liteparse, with a small, hardcoded tolerance — not something this image configures. That check depends on exactly where liteparse thinks each glyph sits, which shifts slightly with the font-rendering stack that produced the page (e.g. a source converted on a different LibreOffice/font stack than this image's). In practice this means: the same source document, converted in two different environments, can land on different sides of that overlap check for a sparse-text page — one keeps clean native text, the other lets a (possibly wrong-language) OCR result overwrite part of it. This showed up concretely testing an Arabic .docx: correct native Arabic text when converted locally on macOS, but a garbled heading when converted inside this image's Debian/LibreOffice, because the container positioned that line's text just outside the overlap tolerance. There's no config knob in liteparse to force "always prefer native text" or to widen the tolerance — low_confidence_pages above is the safety net for this, not a fix; a flagged page means "verify before trusting," not "this is corrected."

⁠Common invoke errors
ErrorCause
source bucket not allowed by guardrail configsource_bucket doesn't match the deployment's ALLOWED_SOURCE_BUCKET
source key does not match required prefixsource_key isn't under the deployment's ALLOWED_SOURCE_PREFIX
destination resolves to the same folder as the source — refusing circular writeDestination folder equals the source's own folder — always rejected, regardless of guardrail config
S3 get object failed: ... NoSuchKey ...source_key doesn't exist — check for a delete marker if the bucket is versioned
missing required field: <field>One of source_bucket/source_key/destination_bucket was omitted

Tag summary

Content type

Image

Digest

sha256:7e701b4c4…

Size

480.2 MB

Last updated

about 2 months ago

docker pull datapsycho/oh-my-liteparser-lambda-v2