Sign inSign up

datapsycho/xpdf-serverless

By datapsycho

•Updated 8 months ago

Serverless PDF conversion service for AWS Lambda.

Image
Machine learning & AI
1

1.6K

datapsycho/xpdf-serverless repository overview

⁠XPDF

⁠XPDF Serverless

Serverless PDF conversion service for AWS Lambda - converts documents (Word, Excel, PowerPoint) to PDF using LibreOffice.

Supports 4 Operation Modes:

  1. S3 Trigger - Automatic conversion on S3 upload
  2. S3-to-S3 - Async conversion with explicit destination
  3. S3-to-Response - Sync conversion returning base64 PDF
  4. Binary-to-Response - Direct binary input returning base64 PDF

⁠Overview

This service converts documents to PDF by:

  1. Detecting S3 file uploads via event notifications (filtered by file type)
  2. Downloading the document to Lambda's /tmp storage
  3. Converting to PDF using LibreOffice CLI (headless mode)
  4. Uploading the PDF back to S3 or returning base64-encoded PDF

Supported formats:

  • Word (.docx, .doc)
  • Excel (.xlsx, .xls)
  • PowerPoint (.pptx, .ppt)

System Specifications:

  • Base OS: Ubuntu 22.04
  • LibreOffice: 7.3.x (headless conversion engine)
  • Python: 3.11
  • Platform: linux/amd64 (x86_64) - required for AWS Lambda
  • Resources: 3008 MB memory (1 full vCPU), 15-minute timeout
  • Current Version: 0.3.7 "Flexible Mode"

⁠Environment Variables

Optional for all modes (can be provided in event payload instead):

  • SOURCE_BUCKET - S3 bucket for input documents (optional constraint)
  • SOURCE_PREFIX - S3 prefix for input (optional constraint, e.g., bronze/)
  • DESTINATION_BUCKET - S3 bucket for output PDFs (optional constraint)
  • DESTINATION_PREFIX - S3 prefix for output (optional constraint, e.g., bronze/pdf/)

Flexible Mode: Environment variables act as optional policy constraints. If set, they validate event-provided values. If not set, the handler uses event-provided bucket/prefix values directly. This enables single Lambda to serve multiple projects without reconfiguration.

⁠Operation Modes

Mode 1: S3 Trigger - Auto-convert on S3 upload
Mode 2: S3-to-S3 - Async conversion with explicit source/destination
Mode 3: S3-to-Response - Sync conversion returning base64 PDF
Mode 4: Binary-to-Response - Direct binary input returning base64 PDF

⁠Payload Examples

Mode 1: S3 Trigger (Auto-triggered by S3 event - no manual invocation needed)

# Automatically triggered when file uploaded to S3
# No JSON payload required - Lambda automatically processes S3 event

Mode 2: S3-to-S3 (Async with explicit destination)

{
  "operation": "s3-to-s3",
  "source_bucket": "my-documents",
  "source_key": "incoming/contract.docx",
  "destination_bucket": "my-pdfs",
  "destination_prefix": "processed"
}

Output: s3://my-pdfs/processed/contract.docx.pdf

Mode 3: S3-to-Response (Sync returning base64 PDF)

{
  "operation": "s3-to-response",
  "source_bucket": "my-documents",
  "source_key": "incoming/report.xlsx",
  "compress": false
}

Returns: JSON with base64-encoded PDF in base64_pdf field

Mode 4: Binary-to-Response (Direct binary input)

{
  "operation": "binary-to-response",
  "filename": "presentation.pptx",
  "file_data": "base64_encoded_binary_content_here...",
  "compress": false
}

Returns: JSON with base64-encoded PDF in base64_pdf field

For detailed payload formats and examples, see docs/progress.md⁠.


⁠Quick Start

Fastest option - No build required. Image is pre-built and tested.

⁠Step 1: Pull the Image
docker pull datapsycho/xpdf-serverless:latest
⁠Step 2: Authenticate with ECR
AWS_ACCOUNT_ID=YOUR-ACCOUNT-ID
AWS_REGION=eu-west-1

aws ecr get-login-password --region $AWS_REGION | \
  docker login --username AWS --password-stdin \
  ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com
⁠Step 3: Tag and Push to ECR
ECR_REPO=xpdf-serverless

docker tag datapsycho/xpdf-serverless:latest \
  ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest

docker push ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest
⁠Step 4: Create or Update Lambda Function

Option A1: Create new function

SOURCE_BUCKET=your-input-bucket
DESTINATION_BUCKET=your-output-bucket
SOURCE_PREFIX=bronze
DESTINATION_PREFIX=bronze

aws lambda create-function \
  --function-name xpdf \
  --role arn:aws:iam::${AWS_ACCOUNT_ID}:role/lambda-execution-role \
  --code ImageUri=${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest \
  --package-type Image \
  --memory-size 3008 \
  --timeout 900 \
  --environment Variables="{SOURCE_BUCKET=${SOURCE_BUCKET},DESTINATION_BUCKET=${DESTINATION_BUCKET},SOURCE_PREFIX=${SOURCE_PREFIX},DESTINATION_PREFIX=${DESTINATION_PREFIX}}" \
  --region ${AWS_REGION}

Option A2: Update existing function

aws lambda update-function-code \
  --function-name xpdf \
  --image-uri ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest \
  --region ${AWS_REGION}
⁠Step 5: Verify Deployment (Optional)
aws lambda get-function \
  --function-name xpdf \
  --region ${AWS_REGION} | jq '{CodeSha256: .Configuration.CodeSha256, LastModified: .Configuration.LastModified, MemorySize: .Configuration.MemorySize}'

The CodeSha256 should match the ECR image digest from Step 3.

⁠Step 6: Test S3 Trigger Mode
# Upload a document to trigger Lambda
aws s3 cp ./document.docx s3://${SOURCE_BUCKET}/${SOURCE_PREFIX}/document.docx \
  --region ${AWS_REGION}

# Monitor logs
aws logs tail /aws/lambda/xpdf --follow --region ${AWS_REGION}

# Verify output in S3
aws s3 ls s3://${DESTINATION_BUCKET}/${DESTINATION_PREFIX}/ --region ${AWS_REGION}

⁠Option B: Build and Deploy Locally
# 1. Build Docker images
docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .
docker build --platform linux/amd64 -t xpdf-serverless:0.3.5 .

# 2. Push to ECR (replace YOUR-ACCOUNT-ID and region as needed)
aws ecr get-login-password --region eu-west-1 | docker login --username AWS --password-stdin YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com
docker tag xpdf-serverless:0.3.5 YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com/xpdf-serverless:0.3.5
docker push YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com/xpdf-serverless:0.3.5

# 3. Deploy Lambda (see AWS Lambda Setup section for IAM role creation)
aws lambda update-function-code \
  --function-name xpdf \
  --image-uri YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com/xpdf-serverless:0.3.5 \
  --region eu-west-1

# 4. Test conversion (S3 Trigger mode)
aws s3 cp test_documents/demo.docx s3://YOUR-BUCKET-NAME/bronze/demo.docx --region eu-west-1
aws s3 ls s3://YOUR-BUCKET-NAME/bronze/  # Check for demo.docx.pdf

⁠Building the Docker Image

This project uses a two-stage build for optimization:

  1. Base image (Dockerfile.base) - Contains LibreOffice and Python (rarely changes)
  2. Service image (Dockerfile) - Adds Lambda runtime and handler code (changes frequently)
⁠Build Process

Step 1: Build base image (required for AWS Lambda x86_64 architecture):

docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .

Step 2: Build service image:

docker build --platform linux/amd64 -t xpdf-serverless:latest .

Why --platform linux/amd64?

  • AWS Lambda requires x86_64 architecture
  • Without this flag, images built on Apple Silicon will be ARM64 and fail with exec format error

Note for Apple Silicon users: You'll see platform mismatch warnings - this is expected and harmless.


⁠Running Locally with Docker

Test the Docker image locally before deploying to AWS Lambda. This is the fastest way to validate conversions.

⁠Quick Test (Single File)
# 1. Pull the image from Docker Hub
docker pull datapsycho/xpdf-serverless:latest

# 2. Create input directory with a test document
mkdir -p test_documents
cp your_document.docx test_documents/

# 3. Run the conversion
docker run -v $(pwd)/test_documents:/tmp/test_documents \
  datapsycho/xpdf-serverless:latest \
  python3 /var/task/src/xpdf/handler.py /tmp/test_documents/your_document.docx

# 4. Check output
ls -lh test_documents/
# You should see: your_document.docx.pdf

Supported Input Formats:

  • .docx, .doc (Word)
  • .xlsx, .xls (Excel)
  • .pptx, .ppt (PowerPoint)
⁠Detailed Example
# Convert demo.docx to demo.pdf
docker run \
  --platform linux/amd64 \
  -v /absolute/path/to/test_documents:/tmp/test_documents \
  datapsycho/xpdf-serverless:latest \
  python3 /var/task/src/xpdf/handler.py /tmp/test_documents/demo.docx

# Output
# ✓ Download successful
# ✓ PDF ready: /tmp/test_documents/demo.pdf
# ✓ Cleaned up: /tmp/test_documents/demo.docx
⁠Parameters Explained
ParameterPurpose
--platform linux/amd64Ensure x86_64 architecture (required for Lambda compatibility)
-v /local/path:/tmp/test_documentsMount local directory with input files; output PDF created here
datapsycho/xpdf-serverless:latestDocker image (pull from Docker Hub)
/var/task/src/xpdf/handler.pyPath inside container (DO NOT CHANGE)
/tmp/test_documents/your_document.docxInput file path (must match mounted volume)
⁠Output Files

After running the command, the converted PDF appears in your local test_documents/ directory:

test_documents/
├── demo.docx           # Input file (unchanged)
└── demo.pdf            # Output file (auto-created)

The PDF filename is automatically generated: {original_filename}.pdf

⁠Notes

File Permissions:

  • Docker runs with user root inside the container
  • Output files may have restrictive permissions on macOS
  • If needed, adjust with: chmod 644 test_documents/*.pdf

macOS Apple Silicon (M1/M2/M3):

WARNING: The requested image's platform (linux/amd64) does not match the detected host platform (linux/arm64/v8)

This warning is expected. The image is built for x86_64 (AWS Lambda requirement), and macOS Docker emulates it transparently. Conversions still work perfectly.

Performance:

  • First run: 10-15 seconds (LibreOffice startup)
  • Subsequent runs: 5-8 seconds (container cache)
  • Large/complex documents: May take longer

⁠Building the Docker Image Locally

If you want to rebuild from source instead of pulling from Docker Hub:

⁠Prerequisites
# Install Docker
# Ensure you're on the latest version
docker --version

# For Apple Silicon, verify platform support
docker run --platform linux/amd64 --rm hello-world
⁠Build Commands

Step 1: Build base image (LibreOffice + Python):

docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .

Step 2: Build service image (adds Lambda runtime + handler):

docker build --platform linux/amd64 -t xpdf-serverless:latest .

Why --platform linux/amd64?

  • AWS Lambda requires x86_64 architecture
  • Without this flag on Apple Silicon, images will be ARM64 and fail with exec format error
⁠Two-Stage Build Process

This project uses a two-stage Docker build for efficiency:

  1. Dockerfile.base (xpdf-python-base)

    • Base: Ubuntu 22.04
    • Installs: LibreOffice 7.3, Python 3.11
    • Size: ~1.9GB
    • Changes rarely → cached
  2. Dockerfile (xpdf-serverless)

    • Builds on: xpdf-python-base
    • Adds: AWS Lambda runtime, handler code
    • Size: ~2.0GB
    • Changes frequently → rebuilt each time
⁠Build Times
  • First build: 15-20 minutes (downloads and installs everything)
  • Subsequent builds: 5-10 seconds (layers cached)
⁠Verify Build
# List images
docker images | grep xpdf

# Test locally
docker run -v $(pwd)/test_documents:/tmp/test_documents \
  xpdf-serverless:latest \
  python3 /var/task/src/xpdf/handler.py /tmp/test_documents/demo.docx

⁠Detailed Setup Guide

For comprehensive AWS setup (IAM roles, ECR, S3 notifications, cleanup), see docs/dev-guide.md⁠. "Id": "pptx-trigger", "LambdaFunctionArn": "arn:aws:lambda:eu-west-1:YOUR-ACCOUNT-ID:function:xpdf", "Events": ["s3:ObjectCreated:*"], "Filter": { "Key": { "FilterRules": [ { "Name": "suffix", "Value": ".pptx" } ] } } } ] } EOF

⁠Apply the notification configuration

aws s3api put-bucket-notification-configuration
--bucket YOUR-BUCKET-NAME
--notification-configuration file://s3-notification.json
--region eu-west-1


**⚠️ Critical:** File type filters prevent infinite loops - without them, the Lambda would trigger on its own PDF outputs!

**Sample Configuration File:**
A template `s3-notification-sample.json` is provided in the repository with placeholders for your specific values.

---

## Testing

Use the integration test suite to validate all 4 operation modes:

```bash
uv run src/tool/integration_check.py

⁠Logging

The Lambda function logs detailed information to CloudWatch:

  • Incoming event - Full S3 event JSON
  • Download progress - File size and path
  • LibreOffice conversion - Command executed, stdout, stderr, return code
  • PDF result - Output file size and upload confirmation
  • Errors - Full exception details with traceback

View logs in real-time:

aws logs tail /aws/lambda/xpdf --follow --region eu-west-1

⁠Configuration

⁠Environment Variables (Set in Dockerfile)

Critical for LibreOffice in Lambda:

  • HOME=/tmp - LibreOffice needs writable home directory (Lambda root filesystem is read-only)
  • SAL_NO_STARTUP_WIZARD=1 - Disables LibreOffice startup wizard

These are required for LibreOffice to function in Lambda's restricted environment.

⁠Timeout & Memory

Production settings:

  • Timeout: 900 seconds (15 minutes) - needed for complex/large documents
  • Memory: 3008 MB - provides 1 full vCPU for faster conversions

Conversion performance example:

  • 1.3 MB Word document → 104 KB PDF in ~88 seconds with 1 vCPU

Smaller values may cause timeouts. Adjust with:

aws lambda update-function-configuration \
  --function-name xpdf-serverless \
  --timeout 900 \
  --memory-size 3008 \
  --region eu-west-1
⁠LibreOffice Flags

Handler uses these flags for conversion:

  • --headless - No GUI
  • --invisible - Completely hidden
  • --nodefault - Don't load default document
  • --nolockcheck - Critical: Prevents filesystem lock issues in Lambda
  • --convert-to pdf - Output format
  • --outdir /tmp - Write to writable directory
⁠Supported Regions

Change eu-west-1 in all commands to your desired region (e.g., us-east-1, us-west-2).


⁠Troubleshooting

⁠"exec format error" in Lambda

Cause: Image built for wrong architecture (ARM64 instead of x86_64)
Fix: Rebuild with --platform linux/amd64 flag:

docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .
docker build --platform linux/amd64 -t xpdf-serverless:0.3.7 .
⁠LibreOffice Timeout Errors

Symptoms: TimeoutExpired error after 45-300 seconds
Causes:

  • Complex/large documents need more time
  • Insufficient CPU allocation

Fixes:

  1. Increase Lambda memory (more memory = more CPU):
    aws lambda update-function-configuration --function-name xpdf --memory-size 3008 --region eu-west-1
    
  2. Increase timeout:
    aws lambda update-function-configuration --function-name xpdf --timeout 900 --region eu-west-1
    
⁠"Read-only file system" Error

Cause: LibreOffice trying to write to home directory
Fix: Ensure ENV HOME=/tmp is set in Dockerfile (already configured)

⁠Function Not Triggering

Check S3 event notification:

aws s3api get-bucket-notification-configuration --bucket YOUR-BUCKET-NAME --region eu-west-1

Ensure filters match your file types (.docx, .xlsx, .pptx)

⁠PDF Not Appearing
  1. Check Lambda execution logs:
    aws logs tail /aws/lambda/xpdf --since 10m --region eu-west-1
    
  2. Verify IAM role has s3:GetObject and s3:PutObject permissions
  3. Check document format is supported (see "Supported formats" above)
⁠Infinite Lambda Invocations

Cause: S3 notification triggering on PDF outputs
Fix: Ensure S3 filters exclude .pdf files (see Step 8 configuration with suffix filters)


⁠Cleanup

See docs/dev-guide.md⁠ for complete AWS resource cleanup instructions.


⁠Architecture

S3 Upload → S3 Event → Lambda Invocation → Download File
                                            ↓
                                    LibreOffice Convert
                                            ↓
                                    Upload PDF to S3

Handler: handler.py⁠
Docker Image: Dockerfile⁠
AI Agent Instructions: .github/copilot-instructions.md⁠

docker run -v $(pwd)/test_documents:/tmp/test_documents libreoffice-service:latest python3 handler.py /tmp/test_documents/file_example_XLSX_10.xlsx

Tag summary

Content type

Image

Digest

sha256:5db3be55b…

Size

751.4 MB

Last updated

8 months ago

docker pull datapsycho/xpdf-serverless