Document-intelligence Lambda container image (Rust). via liteparse.
224
Document-intelligence Lambda container image (Rust). Converts PDF/DOCX/DOC/image
sources to Markdown via liteparse, with
selective OCR (only pages that actually need it), and is invoked directly with an
S3-to-S3 config — no API Gateway/HTTP involved.
Architecture: arm64 only. Base image: public.ecr.aws/lambda/provided:al2023.
arm64. The Lambda function's --architectures must be set to arm64; there is no
x86_64 build. Deploying with --architectures x86_64 (or letting it default) fails
every invoke with Runtime.InvalidEntrypoint/ProcessSpawnFailed — this is an
architecture mismatch, not a bug to work around, and switching to x86_64 is not a
supported alternative for this image.{app_version}-liteparse{liteparse_version}, e.g. 2.0.0-liteparse2.11.1 — always
deploy against an exact tag like this, never latest. latest is published as a
convenience pointer for docker pull only and always tracks the newest release; it is
not a stable, reproducible target and should never be what a live Lambda function points
at.
AWS Lambda only pulls container images from Amazon ECR — not Docker Hub. To deploy this image, pull it here, then push it into your own ECR before pointing Lambda at it:
docker pull datapsycho/oh-my-liteparser-lambda-v2:<TAG>
AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
AWS_REGION=<your-region>
ECR_URI="${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/oh-my-liteparser-lambda-v2"
aws ecr create-repository --repository-name oh-my-liteparser-lambda-v2 --region "${AWS_REGION}" # first time only
docker tag datapsycho/oh-my-liteparser-lambda-v2:<TAG> "${ECR_URI}:<TAG>"
aws ecr get-login-password --region "${AWS_REGION}" \
| docker login --username AWS --password-stdin "${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com"
docker push "${ECR_URI}:<TAG>"
Trusted for lambda.amazonaws.com, with exactly the S3 access this function needs —
scope Resource to your actual bucket(s), not *:
{
"Version": "2012-10-17",
"Statement": [
{ "Sid": "ReadSourceDocuments", "Effect": "Allow", "Action": ["s3:GetObject", "s3:HeadObject"], "Resource": "arn:aws:s3:::<your-bucket>/*" },
{ "Sid": "WriteDestinationOutput", "Effect": "Allow", "Action": ["s3:PutObject"], "Resource": "arn:aws:s3:::<your-bucket>/*" },
{ "Sid": "ListBucket", "Effect": "Allow", "Action": ["s3:ListBucket", "s3:GetBucketLocation"], "Resource": "arn:aws:s3:::<your-bucket>" }
]
}
Plus the AWS-managed AWSLambdaBasicExecutionRole policy, for CloudWatch Logs.
--architectures arm64 must match this image's actual architecture — mismatching it
fails every invoke with Runtime.InvalidEntrypoint/ProcessSpawnFailed.
aws lambda create-function \
--function-name <your-function-name> \
--package-type Image \
--code ImageUri="${ECR_URI}:<TAG>" \
--role arn:aws:iam::${AWS_ACCOUNT_ID}:role/<your-execution-role> \
--timeout 300 \
--memory-size 3008 \
--architectures arm64 \
--region "${AWS_REGION}"
300s timeout / 3008MB memory are verified comfortable for real documents up to 362 pages
with mixed OCR (~107s duration, ~306MB peak memory). Raise either later via
update-function-configuration if your documents are larger or more OCR-heavy — no image
rebuild needed. /tmp (ephemeral storage) defaults to 512MB; raise it
(--ephemeral-storage Size=<MB>) if you expect very large or image-heavy source
documents.
| Variable | Example | Purpose |
|---|---|---|
ALLOWED_SOURCE_BUCKET | my-bucket | Reject any invoke naming a different source bucket |
ALLOWED_SOURCE_PREFIX | input/ | Reject any invoke whose source key isn't under this prefix |
ALLOWED_DEST_BUCKET | my-bucket | Reject any invoke naming a different destination bucket |
ALLOWED_DEST_PREFIX | output/ | Reject any invoke whose destination isn't under this prefix |
AWS_LAMBDA_LOG_FORMAT | JSON | Structured JSON log lines (queryable in CloudWatch Logs Insights) |
AWS_LAMBDA_LOG_LEVEL / RUST_LOG | INFO / info | Log verbosity |
These guardrails are deploy-time only — not part of the invoke payload, so a caller can't loosen them. A request whose destination folder would equal its own source folder is always rejected regardless of these settings (circular-write protection).
aws lambda update-function-configuration \
--function-name <your-function-name> \
--region "${AWS_REGION}" \
--environment 'Variables={ALLOWED_SOURCE_BUCKET=<bucket>,ALLOWED_DEST_BUCKET=<bucket>,ALLOWED_SOURCE_PREFIX=input/,ALLOWED_DEST_PREFIX=output/,AWS_LAMBDA_LOG_FORMAT=JSON,RUST_LOG=info}' \
--logging-config LogFormat=JSON,ApplicationLogLevel=INFO,SystemLogLevel=INFO
Repoint at a new tag — nothing else changes:
aws lambda update-function-code \
--function-name <your-function-name> \
--image-uri "${ECR_URI}:<NEW_TAG>" \
--region "${AWS_REGION}"
The caller (a service, a scheduled job, a human via aws lambda invoke) needs only
lambda:InvokeFunction on this one function's ARN — no S3 permissions at all; the
execution role above handles all S3 access on the caller's behalf. Ask your DevOps/infra
team to attach this to your IAM user/role:
{
"Version": "2012-10-17",
"Statement": [
{ "Sid": "InvokeDocumentConversion", "Effect": "Allow", "Action": "lambda:InvokeFunction", "Resource": "arn:aws:lambda:<region>:<account-id>:function:<your-function-name>" }
]
}
Scoped to that one function ARN, not function:*, so it can't be used to invoke anything
else in the account.
{
"source_bucket": "my-input-bucket",
"source_key": "input/report.docx",
"destination_bucket": "my-output-bucket",
"destination_prefix": "output/", // optional, default ""
"destination_key": null, // optional — overrides prefix + derived name
"image_mode": "placeholder", // optional — off | placeholder | embed
"extract_links": true, // optional, default true
"ocr_mode": "auto", // optional — auto | always | never
"ocr_language": null, // optional — see "OCR languages" below
"password": null, // optional — encrypted PDF/Office sources
"plain": false // optional — skip front matter/page markers
}
Only source_bucket/source_key/destination_bucket are required; everything else has
a default. source_bucket/destination_bucket (and prefixes, if ALLOWED_* guardrails
are set) must satisfy whatever the deploying account configured — ask DevOps which
bucket/prefix you're allowed to use.
The image ships Tesseract trained data for English plus the common Western European
Latin-script languages: eng, fra, deu, spa, ita, por, nld (CJK is out of
scope for this OCR engine). Leave ocr_language unset and all of these are recognized
together in a single pass — Tesseract loads every dictionary at once and picks the
best-matching one per word/region, so you don't need to know or specify a language up
front. Set ocr_language to a single code (or your own +-joined subset, e.g.
"fra+deu") to force specific language(s) — typically only useful for a small speed gain
when you already know the document's language. An unrecognized code, or one without
matching trained data in the image, fails the OCR step at invoke time.
This limitation only applies when OCR actually runs. Whether a page needs OCR at all is decided purely by text density/images/cmap sanity — never by script or language. A Japanese, Arabic, or Russian source with a real embedded text layer (not a scan) skips OCR entirely and its text is extracted as-is, so the Latin-only tessdata above is irrelevant to it. The limitation only bites a scanned/image-only non-Latin page, where OCR is the only way to get any text — Tesseract still produces output there, but confidently misreads non-Latin glyphs as something Latin, which is what the low-confidence signal below is for.
aws lambda invoke \
--function-name <your-function-name> \
--region <region> \
--cli-binary-format raw-in-base64-out \
--payload '{"source_bucket":"<bucket>","source_key":"input/report.docx","destination_bucket":"<bucket>","destination_prefix":"output/"}' \
response.json && cat response.json
{
"source": { "bucket": "...", "key": "..." },
"destination": { "bucket": "...", "key": "..." },
"pages": 5,
"ocr_enabled": true,
"ocr_engine": "tesseract",
"ocr_pages_used": 3,
"low_confidence_pages": [],
"average_ocr_confidence": 0.87,
"detected_language": "fra",
"detected_language_name": "French",
"detected_script": "Latin",
"detected_language_confidence": 1.0,
"detected_language_reliable": true
}
ocr_pages_used counts pages OCR was attempted on, not pages whose text actually
came from OCR. The engine discards any OCR result that overlaps text a page already
has natively, so a page can count toward ocr_pages_used while the Markdown that lands
in S3 is still 100% native text for that page — check the per-page
ocr_contributed_text marker (below) to know which pages' text genuinely came from OCR.
average_ocr_confidence is a document-level mean of each OCR'd page's own
confidence — a quick glance, not a substitute for low_confidence_pages: a single
badly-OCR'd page among many clean ones can still average out fine, so the per-page list
stays the actual diagnostic. null when no page was OCR'd.
detected_language/_name/_script/_confidence/_reliable come from running
whatlang over a bounded sample of the extracted
text (first ~5 pages / ~3000 chars) — pure metadata with no effect on
ocr_language/OCR routing, which stays entirely separate and under caller control.
detected_language/_name is which language (French, English, ...);
detected_script is which writing system (Latin, Cyrillic, Arabic, ...) — many
unrelated languages share a script, so detected_script alone can't distinguish English
from French (both "Latin"), but it's a fast way to tell whether a document is even in
a script this image's Latin-only tessdata can OCR at all. detected_language_reliable == false is still a real prediction, just low-confidence — surfaced, not hidden. All five
detected_* fields are null together only when detection found nothing usable.
The converted Markdown itself lands at destination.bucket/destination.key in S3 —
this response is only the summary, not the content. Markdown includes a YAML front-matter
block (source, filename, pages, ocr_enabled, ocr_engine, ocr_pages_used,
low_confidence_pages, average_ocr_confidence, the five detected_* fields above,
parsed_at) followed by per-page content with a marker — a JSON
object with a fixed key set on every page (ocr_confidence is null, not omitted, when
there's nothing to report), so it can be parsed with any JSON library rather than
regex/string-split:
<!-- page: {"page": N, "ocr": bool, "ocr_contributed_text": bool, "ocr_confidence": number | null} -->
ocr | ocr_contributed_text | ocr_confidence | Meaning |
|---|---|---|---|
false | false | null | No OCR attempted — page didn't need it. |
true | false | null | OCR attempted, but every result overlapped existing native text and was discarded — the page is still 100% native text. ocr: true here does not mean the text you're reading came from OCR. |
true | true | 0.61 | OCR attempted and its text survived into the page — ocr_confidence is the average Tesseract confidence across just that OCR-sourced text. |
(unless "plain": true was set, which omits front matter and markers entirely).
low_confidence_pages lists page numbers where ocr_contributed_text was true and the
average confidence scored below the threshold (0.75) — a quality signal, not a
language detector. It commonly catches OCR run against the wrong ocr_language (e.g. a
non-Latin-script scan, which this image can't recognize at all — see "OCR languages"
above), but a genuinely poor scan in the correct language can trip it too. Treat a
flagged page as "don't trust this text blindly," not as a hard failure — the request
still succeeds.
Caveat: the native-vs-OCR decision can differ across environments for the same input.
Whether an OCR result is discarded (kept as native text) or kept (merged/overrides that
region) comes down to a bounding-box overlap check inside liteparse, with a small,
hardcoded tolerance — not something this image configures. That check depends on
exactly where liteparse thinks each glyph sits, which shifts slightly with the
font-rendering stack that produced the page (e.g. a source converted on a different
LibreOffice/font stack than this image's). In practice this means: the same source
document, converted in two different environments, can land on different sides of that
overlap check for a sparse-text page — one keeps clean native text, the other lets a
(possibly wrong-language) OCR result overwrite part of it. This showed up concretely
testing an Arabic .docx: correct native Arabic text when converted locally on macOS,
but a garbled heading when converted inside this image's Debian/LibreOffice, because the
container positioned that line's text just outside the overlap tolerance. There's no
config knob in liteparse to force "always prefer native text" or to widen the
tolerance — low_confidence_pages above is the safety net for this, not a fix; a
flagged page means "verify before trusting," not "this is corrected."
| Error | Cause |
|---|---|
source bucket not allowed by guardrail config | source_bucket doesn't match the deployment's ALLOWED_SOURCE_BUCKET |
source key does not match required prefix | source_key isn't under the deployment's ALLOWED_SOURCE_PREFIX |
destination resolves to the same folder as the source — refusing circular write | Destination folder equals the source's own folder — always rejected, regardless of guardrail config |
S3 get object failed: ... NoSuchKey ... | source_key doesn't exist — check for a delete marker if the bucket is versioned |
missing required field: <field> | One of source_bucket/source_key/destination_bucket was omitted |
Content type
Image
Digest
sha256:7e701b4c4…
Size
480.2 MB
Last updated
about 2 months ago
docker pull datapsycho/oh-my-liteparser-lambda-v2