AWS Lambda container image for serverless document parsing with liteparse
207
AWS Lambda container image for serverless document parsing. Reads source files from S3, parses them using LiteParse, and writes structured output back to S3. Built on top of the oh-my-liteparser base image.
document-parsing · aws-lambda · serverless · s3 · liteparse · ocr · nodejs
datapsycho/oh-my-liteparser-lambda is the AWS Lambda container image for the oh-my-liteparser project. It extends the batteries-included base image (datapsycho/oh-my-liteparser) with:
dist/lambda/handler.js)aws-lambda-ric) as the entrypoint@aws-sdk/client-s3) for storage orchestrationIt is designed to be deployed as an AWS Lambda function using container image support. When invoked, it reads source documents from S3, runs them through the same LiteParse parsing core used by the CLI image, and writes the parsed results back to a specified S3 location.
Running document parsing as a Lambda function normally requires managing a deployment package with all native dependencies — LibreOffice, ImageMagick, Ghostscript, Tesseract — which easily exceed Lambda's unzipped size limits and are difficult to layer correctly.
This image solves that by using Lambda's container image support. The full parsing runtime ships as a self-contained Docker image, bypassing zip size constraints entirely. Deploying becomes a matter of pushing the image to ECR and pointing a Lambda function at it.
This image builds on datapsycho/oh-my-liteparser and adds:
dist/lambda/handler.js)aws-lambda-ric — AWS Lambda Runtime Interface Client (entrypoint)@aws-sdk/client-s3 — S3 read/write orchestrationtslib — TypeScript runtime helpersAll system-level dependencies (LibreOffice, ImageMagick, Ghostscript, Tesseract, LiteParse) are inherited from the base image.
The Lambda function handler is:
dist/lambda/handler.handler
Direct invoke payload:
{
"source_bucket": "my-input-bucket",
"source_prefix": "input/",
"source_key": "documents/report.pdf",
"destination_bucket": "my-output-bucket",
"destination_prefix": "parsed/",
"format": "md"
}
In direct invoke mode, the source document must already exist in S3. The payload tells the Lambda where to download the source object from and where to upload the rendered output.
The Lambda also supports native S3 trigger invocations from bucket notifications under the configured source prefix. In both modes, the handler:
/tmpCurrent behavior is intentionally split this way:
OMLP_SOURCE_PREFIX auto-invoke the LambdaOMLP_DESTINATION_PREFIX, but direct invokes can override them in the payloadTrigger example:
s3://ohmyliteparse/input/docx/demo.docxs3://ohmyliteparse/output/docx/demo.docx.md.jsonDirect invoke example:
{
"source_bucket": "ohmyliteparse",
"source_prefix": "input/",
"source_key": "nested/report.pdf",
"destination_bucket": "ohmyliteparse",
"destination_prefix": "output/",
"format": "liteparse-json",
"ocr": false
}
s3://ohmyliteparse/input/nested/report.pdfs3://ohmyliteparse/input/nested/report.pdfs3://ohmyliteparse/output/nested/report.pdf.liteparse.jsonBuild and push to ECR:
# Authenticate with ECR
aws ecr get-login-password --region us-east-1 | \
docker login --username AWS --password-stdin <account-id>.dkr.ecr.us-east-1.amazonaws.com
# Build
docker build \
--platform linux/amd64 \
-f docker/Dockerfile.lambda \
-t oh-my-liteparser-lambda:1.5.2 .
# Tag and push
docker tag oh-my-liteparser-lambda:1.5.2 \
<account-id>.dkr.ecr.us-east-1.amazonaws.com/oh-my-liteparser-lambda:1.5.2
docker push <account-id>.dkr.ecr.us-east-1.amazonaws.com/oh-my-liteparser-lambda:1.5.2
For this repository's standard AWS deployment flow, prefer:
bash infra/scripts/01-build-push-ecr.sh
bash infra/scripts/02-deploy-cfn.sh
Create the Lambda function (AWS CLI):
aws lambda create-function \
--function-name oh-my-liteparser \
--package-type Image \
--code ImageUri=<account-id>.dkr.ecr.us-east-1.amazonaws.com/oh-my-liteparser-lambda:1.5.2 \
--role arn:aws:iam::<account-id>:role/lambda-execution-role \
--timeout 300 \
--memory-size 3008
Required IAM permissions for the Lambda execution role:
s3:GetObject and s3:HeadObject on objects in the buckets3:PutObject on objects in the buckets3:ListBucket and s3:GetBucketLocation on the bucketThis means:
PDF, DOCX, PPTX, XLSX, ODT, RTF, JPG, PNG, TIFF, WEBP, SVG, BMP, GIF — any format supported by LiteParse.
| Format value | Description |
|---|---|
plain | Text-oriented markdown written as .txt.md |
md | Rich per-page markdown envelope with both pages[] and merged content |
liteparse-json | Full raw LiteParse result with metadata |
screenshot | Screenshot bundle with manifest, per-page PNG, and per-page text |
The deployed Lambda commonly relies on these environment variables:
| Variable | Purpose |
|---|---|
OMLP_SOURCE_BUCKET | Default source bucket |
OMLP_SOURCE_PREFIX | Source prefix watched by the S3 trigger |
OMLP_DESTINATION_BUCKET | Default destination bucket |
OMLP_DESTINATION_PREFIX | Destination prefix for rendered output |
OMLP_FORMAT | Default output format when the event omits format |
OMLP_OCR | Default OCR mode (true or false) |
OMLP_CONFIDENCE | Optional default OCR confidence threshold |
| Tag | Description |
|---|---|
1.5.2 | Pinned release aligned with LiteParse v1.5.x |
linux/amd64 only. Required for AWS Lambda container runtime compatibility.
The document parsing capability in this image is powered by LiteParse, developed by the LlamaIndex team. LiteParse provides structured, page-aware parsing across a wide range of document formats with optional OCR support. All credit for the underlying parsing engine goes to the LlamaIndex LiteParse contributors.
This Lambda image wraps LiteParse with S3 orchestration and a Lambda-compatible handler entrypoint, enabling serverless document parsing workflows without the complexity of managing native dependencies in a traditional Lambda deployment package.
Content type
Image
Digest
sha256:b88acd8b3…
Size
515.1 MB
Last updated
5 months ago
docker pull datapsycho/oh-my-liteparser-lambda