Sign inSign up

datapsycho/oh-my-liteparser-lambda

By datapsycho

•Updated 5 months ago

AWS Lambda container image for serverless document parsing with liteparse

Image
Machine learning & AI
Data science
Web servers
0

207

datapsycho/oh-my-liteparser-lambda repository overview

⁠oh-my-liteparser-lambda — Lambda Image

⁠Description

AWS Lambda container image for serverless document parsing. Reads source files from S3, parses them using LiteParse, and writes structured output back to S3. Built on top of the oh-my-liteparser base image.

⁠Categories

document-parsing · aws-lambda · serverless · s3 · liteparse · ocr · nodejs


⁠Repository Overview

⁠What This Image Is

datapsycho/oh-my-liteparser-lambda is the AWS Lambda container image for the oh-my-liteparser project. It extends the batteries-included base image (datapsycho/oh-my-liteparser) with:

  • The compiled Lambda handler (dist/lambda/handler.js)
  • The AWS Lambda Runtime Interface Client (aws-lambda-ric) as the entrypoint
  • The AWS SDK S3 client (@aws-sdk/client-s3) for storage orchestration

It is designed to be deployed as an AWS Lambda function using container image support. When invoked, it reads source documents from S3, runs them through the same LiteParse parsing core used by the CLI image, and writes the parsed results back to a specified S3 location.

⁠What Problem It Solves

Running document parsing as a Lambda function normally requires managing a deployment package with all native dependencies — LibreOffice, ImageMagick, Ghostscript, Tesseract — which easily exceed Lambda's unzipped size limits and are difficult to layer correctly.

This image solves that by using Lambda's container image support. The full parsing runtime ships as a self-contained Docker image, bypassing zip size constraints entirely. Deploying becomes a matter of pushing the image to ECR and pointing a Lambda function at it.

⁠What Is Included

This image builds on datapsycho/oh-my-liteparser and adds:

  • The compiled Lambda handler (dist/lambda/handler.js)
  • aws-lambda-ric — AWS Lambda Runtime Interface Client (entrypoint)
  • @aws-sdk/client-s3 — S3 read/write orchestration
  • tslib — TypeScript runtime helpers

All system-level dependencies (LibreOffice, ImageMagick, Ghostscript, Tesseract, LiteParse) are inherited from the base image.

⁠Handler Invocation

The Lambda function handler is:

dist/lambda/handler.handler

Direct invoke payload:

{
  "source_bucket": "my-input-bucket",
  "source_prefix": "input/",
  "source_key": "documents/report.pdf",
  "destination_bucket": "my-output-bucket",
  "destination_prefix": "parsed/",
  "format": "md"
}

In direct invoke mode, the source document must already exist in S3. The payload tells the Lambda where to download the source object from and where to upload the rendered output.

The Lambda also supports native S3 trigger invocations from bucket notifications under the configured source prefix. In both modes, the handler:

  1. Downloads the source file from S3 into /tmp
  2. Runs the LiteParse parsing flow
  3. Writes the parsed output file(s) to the destination S3 prefix
  4. Returns metadata about the operation

Current behavior is intentionally split this way:

  • trigger scope stays narrow: only objects created under the configured OMLP_SOURCE_PREFIX auto-invoke the Lambda
  • manual invocation stays flexible: direct invokes can point at any object path in the same bucket, not only the trigger prefix
  • destination defaults still come from OMLP_DESTINATION_PREFIX, but direct invokes can override them in the payload

Trigger example:

  • Uploaded object: s3://ohmyliteparse/input/docx/demo.docx
  • Generated object: s3://ohmyliteparse/output/docx/demo.docx.md.json

Direct invoke example:

{
  "source_bucket": "ohmyliteparse",
  "source_prefix": "input/",
  "source_key": "nested/report.pdf",
  "destination_bucket": "ohmyliteparse",
  "destination_prefix": "output/",
  "format": "liteparse-json",
  "ocr": false
}
  • Upload source file first: s3://ohmyliteparse/input/nested/report.pdf
  • Source object: s3://ohmyliteparse/input/nested/report.pdf
  • Generated object: s3://ohmyliteparse/output/nested/report.pdf.liteparse.json
⁠How to Deploy

Build and push to ECR:

# Authenticate with ECR
aws ecr get-login-password --region us-east-1 | \
  docker login --username AWS --password-stdin <account-id>.dkr.ecr.us-east-1.amazonaws.com

# Build
docker build \
  --platform linux/amd64 \
  -f docker/Dockerfile.lambda \
  -t oh-my-liteparser-lambda:1.5.2 .

# Tag and push
docker tag oh-my-liteparser-lambda:1.5.2 \
  <account-id>.dkr.ecr.us-east-1.amazonaws.com/oh-my-liteparser-lambda:1.5.2

docker push <account-id>.dkr.ecr.us-east-1.amazonaws.com/oh-my-liteparser-lambda:1.5.2

For this repository's standard AWS deployment flow, prefer:

bash infra/scripts/01-build-push-ecr.sh
bash infra/scripts/02-deploy-cfn.sh

Create the Lambda function (AWS CLI):

aws lambda create-function \
  --function-name oh-my-liteparser \
  --package-type Image \
  --code ImageUri=<account-id>.dkr.ecr.us-east-1.amazonaws.com/oh-my-liteparser-lambda:1.5.2 \
  --role arn:aws:iam::<account-id>:role/lambda-execution-role \
  --timeout 300 \
  --memory-size 3008

Required IAM permissions for the Lambda execution role:

  • s3:GetObject and s3:HeadObject on objects in the bucket
  • s3:PutObject on objects in the bucket
  • s3:ListBucket and s3:GetBucketLocation on the bucket

This means:

  • the trigger configuration controls which uploads auto-run the Lambda
  • the IAM policy allows direct invokes to read or write other object paths in the same bucket without hitting a contract mismatch
⁠Supported Input Formats

PDF, DOCX, PPTX, XLSX, ODT, RTF, JPG, PNG, TIFF, WEBP, SVG, BMP, GIF — any format supported by LiteParse.

⁠Supported Output Formats
Format valueDescription
plainText-oriented markdown written as .txt.md
mdRich per-page markdown envelope with both pages[] and merged content
liteparse-jsonFull raw LiteParse result with metadata
screenshotScreenshot bundle with manifest, per-page PNG, and per-page text
⁠Default Runtime Environment

The deployed Lambda commonly relies on these environment variables:

VariablePurpose
OMLP_SOURCE_BUCKETDefault source bucket
OMLP_SOURCE_PREFIXSource prefix watched by the S3 trigger
OMLP_DESTINATION_BUCKETDefault destination bucket
OMLP_DESTINATION_PREFIXDestination prefix for rendered output
OMLP_FORMATDefault output format when the event omits format
OMLP_OCRDefault OCR mode (true or false)
OMLP_CONFIDENCEOptional default OCR confidence threshold
⁠Tags
TagDescription
1.5.2Pinned release aligned with LiteParse v1.5.x
⁠Platform

linux/amd64 only. Required for AWS Lambda container runtime compatibility.

⁠Credits

The document parsing capability in this image is powered by LiteParse⁠, developed by the LlamaIndex team. LiteParse provides structured, page-aware parsing across a wide range of document formats with optional OCR support. All credit for the underlying parsing engine goes to the LlamaIndex LiteParse contributors.

This Lambda image wraps LiteParse with S3 orchestration and a Lambda-compatible handler entrypoint, enabling serverless document parsing workflows without the complexity of managing native dependencies in a traditional Lambda deployment package.

Tag summary

Content type

Image

Digest

sha256:b88acd8b3…

Size

515.1 MB

Last updated

5 months ago

docker pull datapsycho/oh-my-liteparser-lambda