Serverless PDF conversion service for AWS Lambda.
1.6K
Serverless PDF conversion service for AWS Lambda - converts documents (Word, Excel, PowerPoint) to PDF using LibreOffice.
Supports 4 Operation Modes:
This service converts documents to PDF by:
/tmp storageSupported formats:
System Specifications:
Optional for all modes (can be provided in event payload instead):
SOURCE_BUCKET - S3 bucket for input documents (optional constraint)SOURCE_PREFIX - S3 prefix for input (optional constraint, e.g., bronze/)DESTINATION_BUCKET - S3 bucket for output PDFs (optional constraint)DESTINATION_PREFIX - S3 prefix for output (optional constraint, e.g., bronze/pdf/)Flexible Mode: Environment variables act as optional policy constraints. If set, they validate event-provided values. If not set, the handler uses event-provided bucket/prefix values directly. This enables single Lambda to serve multiple projects without reconfiguration.
Mode 1: S3 Trigger - Auto-convert on S3 upload
Mode 2: S3-to-S3 - Async conversion with explicit source/destination
Mode 3: S3-to-Response - Sync conversion returning base64 PDF
Mode 4: Binary-to-Response - Direct binary input returning base64 PDF
Mode 1: S3 Trigger (Auto-triggered by S3 event - no manual invocation needed)
# Automatically triggered when file uploaded to S3
# No JSON payload required - Lambda automatically processes S3 event
Mode 2: S3-to-S3 (Async with explicit destination)
{
"operation": "s3-to-s3",
"source_bucket": "my-documents",
"source_key": "incoming/contract.docx",
"destination_bucket": "my-pdfs",
"destination_prefix": "processed"
}
Output: s3://my-pdfs/processed/contract.docx.pdf
Mode 3: S3-to-Response (Sync returning base64 PDF)
{
"operation": "s3-to-response",
"source_bucket": "my-documents",
"source_key": "incoming/report.xlsx",
"compress": false
}
Returns: JSON with base64-encoded PDF in base64_pdf field
Mode 4: Binary-to-Response (Direct binary input)
{
"operation": "binary-to-response",
"filename": "presentation.pptx",
"file_data": "base64_encoded_binary_content_here...",
"compress": false
}
Returns: JSON with base64-encoded PDF in base64_pdf field
For detailed payload formats and examples, see docs/progress.md.
Fastest option - No build required. Image is pre-built and tested.
docker pull datapsycho/xpdf-serverless:latest
AWS_ACCOUNT_ID=YOUR-ACCOUNT-ID
AWS_REGION=eu-west-1
aws ecr get-login-password --region $AWS_REGION | \
docker login --username AWS --password-stdin \
${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com
ECR_REPO=xpdf-serverless
docker tag datapsycho/xpdf-serverless:latest \
${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest
docker push ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest
Option A1: Create new function
SOURCE_BUCKET=your-input-bucket
DESTINATION_BUCKET=your-output-bucket
SOURCE_PREFIX=bronze
DESTINATION_PREFIX=bronze
aws lambda create-function \
--function-name xpdf \
--role arn:aws:iam::${AWS_ACCOUNT_ID}:role/lambda-execution-role \
--code ImageUri=${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest \
--package-type Image \
--memory-size 3008 \
--timeout 900 \
--environment Variables="{SOURCE_BUCKET=${SOURCE_BUCKET},DESTINATION_BUCKET=${DESTINATION_BUCKET},SOURCE_PREFIX=${SOURCE_PREFIX},DESTINATION_PREFIX=${DESTINATION_PREFIX}}" \
--region ${AWS_REGION}
Option A2: Update existing function
aws lambda update-function-code \
--function-name xpdf \
--image-uri ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/${ECR_REPO}:latest \
--region ${AWS_REGION}
aws lambda get-function \
--function-name xpdf \
--region ${AWS_REGION} | jq '{CodeSha256: .Configuration.CodeSha256, LastModified: .Configuration.LastModified, MemorySize: .Configuration.MemorySize}'
The CodeSha256 should match the ECR image digest from Step 3.
# Upload a document to trigger Lambda
aws s3 cp ./document.docx s3://${SOURCE_BUCKET}/${SOURCE_PREFIX}/document.docx \
--region ${AWS_REGION}
# Monitor logs
aws logs tail /aws/lambda/xpdf --follow --region ${AWS_REGION}
# Verify output in S3
aws s3 ls s3://${DESTINATION_BUCKET}/${DESTINATION_PREFIX}/ --region ${AWS_REGION}
# 1. Build Docker images
docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .
docker build --platform linux/amd64 -t xpdf-serverless:0.3.5 .
# 2. Push to ECR (replace YOUR-ACCOUNT-ID and region as needed)
aws ecr get-login-password --region eu-west-1 | docker login --username AWS --password-stdin YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com
docker tag xpdf-serverless:0.3.5 YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com/xpdf-serverless:0.3.5
docker push YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com/xpdf-serverless:0.3.5
# 3. Deploy Lambda (see AWS Lambda Setup section for IAM role creation)
aws lambda update-function-code \
--function-name xpdf \
--image-uri YOUR-ACCOUNT-ID.dkr.ecr.eu-west-1.amazonaws.com/xpdf-serverless:0.3.5 \
--region eu-west-1
# 4. Test conversion (S3 Trigger mode)
aws s3 cp test_documents/demo.docx s3://YOUR-BUCKET-NAME/bronze/demo.docx --region eu-west-1
aws s3 ls s3://YOUR-BUCKET-NAME/bronze/ # Check for demo.docx.pdf
This project uses a two-stage build for optimization:
Dockerfile.base) - Contains LibreOffice and Python (rarely changes)Dockerfile) - Adds Lambda runtime and handler code (changes frequently)Step 1: Build base image (required for AWS Lambda x86_64 architecture):
docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .
Step 2: Build service image:
docker build --platform linux/amd64 -t xpdf-serverless:latest .
Why --platform linux/amd64?
exec format errorNote for Apple Silicon users: You'll see platform mismatch warnings - this is expected and harmless.
Test the Docker image locally before deploying to AWS Lambda. This is the fastest way to validate conversions.
# 1. Pull the image from Docker Hub
docker pull datapsycho/xpdf-serverless:latest
# 2. Create input directory with a test document
mkdir -p test_documents
cp your_document.docx test_documents/
# 3. Run the conversion
docker run -v $(pwd)/test_documents:/tmp/test_documents \
datapsycho/xpdf-serverless:latest \
python3 /var/task/src/xpdf/handler.py /tmp/test_documents/your_document.docx
# 4. Check output
ls -lh test_documents/
# You should see: your_document.docx.pdf
Supported Input Formats:
.docx, .doc (Word).xlsx, .xls (Excel).pptx, .ppt (PowerPoint)# Convert demo.docx to demo.pdf
docker run \
--platform linux/amd64 \
-v /absolute/path/to/test_documents:/tmp/test_documents \
datapsycho/xpdf-serverless:latest \
python3 /var/task/src/xpdf/handler.py /tmp/test_documents/demo.docx
# Output
# ✓ Download successful
# ✓ PDF ready: /tmp/test_documents/demo.pdf
# ✓ Cleaned up: /tmp/test_documents/demo.docx
| Parameter | Purpose |
|---|---|
--platform linux/amd64 | Ensure x86_64 architecture (required for Lambda compatibility) |
-v /local/path:/tmp/test_documents | Mount local directory with input files; output PDF created here |
datapsycho/xpdf-serverless:latest | Docker image (pull from Docker Hub) |
/var/task/src/xpdf/handler.py | Path inside container (DO NOT CHANGE) |
/tmp/test_documents/your_document.docx | Input file path (must match mounted volume) |
After running the command, the converted PDF appears in your local test_documents/ directory:
test_documents/
├── demo.docx # Input file (unchanged)
└── demo.pdf # Output file (auto-created)
The PDF filename is automatically generated: {original_filename}.pdf
File Permissions:
root inside the containerchmod 644 test_documents/*.pdfmacOS Apple Silicon (M1/M2/M3):
WARNING: The requested image's platform (linux/amd64) does not match the detected host platform (linux/arm64/v8)
This warning is expected. The image is built for x86_64 (AWS Lambda requirement), and macOS Docker emulates it transparently. Conversions still work perfectly.
Performance:
If you want to rebuild from source instead of pulling from Docker Hub:
# Install Docker
# Ensure you're on the latest version
docker --version
# For Apple Silicon, verify platform support
docker run --platform linux/amd64 --rm hello-world
Step 1: Build base image (LibreOffice + Python):
docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .
Step 2: Build service image (adds Lambda runtime + handler):
docker build --platform linux/amd64 -t xpdf-serverless:latest .
Why --platform linux/amd64?
exec format errorThis project uses a two-stage Docker build for efficiency:
Dockerfile.base (xpdf-python-base)
Dockerfile (xpdf-serverless)
xpdf-python-base# List images
docker images | grep xpdf
# Test locally
docker run -v $(pwd)/test_documents:/tmp/test_documents \
xpdf-serverless:latest \
python3 /var/task/src/xpdf/handler.py /tmp/test_documents/demo.docx
For comprehensive AWS setup (IAM roles, ECR, S3 notifications, cleanup), see docs/dev-guide.md. "Id": "pptx-trigger", "LambdaFunctionArn": "arn:aws:lambda:eu-west-1:YOUR-ACCOUNT-ID:function:xpdf", "Events": ["s3:ObjectCreated:*"], "Filter": { "Key": { "FilterRules": [ { "Name": "suffix", "Value": ".pptx" } ] } } } ] } EOF
aws s3api put-bucket-notification-configuration
--bucket YOUR-BUCKET-NAME
--notification-configuration file://s3-notification.json
--region eu-west-1
**⚠️ Critical:** File type filters prevent infinite loops - without them, the Lambda would trigger on its own PDF outputs!
**Sample Configuration File:**
A template `s3-notification-sample.json` is provided in the repository with placeholders for your specific values.
---
## Testing
Use the integration test suite to validate all 4 operation modes:
```bash
uv run src/tool/integration_check.py
The Lambda function logs detailed information to CloudWatch:
View logs in real-time:
aws logs tail /aws/lambda/xpdf --follow --region eu-west-1
Critical for LibreOffice in Lambda:
HOME=/tmp - LibreOffice needs writable home directory (Lambda root filesystem is read-only)SAL_NO_STARTUP_WIZARD=1 - Disables LibreOffice startup wizardThese are required for LibreOffice to function in Lambda's restricted environment.
Production settings:
Conversion performance example:
Smaller values may cause timeouts. Adjust with:
aws lambda update-function-configuration \
--function-name xpdf-serverless \
--timeout 900 \
--memory-size 3008 \
--region eu-west-1
Handler uses these flags for conversion:
--headless - No GUI--invisible - Completely hidden--nodefault - Don't load default document--nolockcheck - Critical: Prevents filesystem lock issues in Lambda--convert-to pdf - Output format--outdir /tmp - Write to writable directoryChange eu-west-1 in all commands to your desired region (e.g., us-east-1, us-west-2).
Cause: Image built for wrong architecture (ARM64 instead of x86_64)
Fix: Rebuild with --platform linux/amd64 flag:
docker build --platform linux/amd64 -f Dockerfile.base -t xpdf-python-base:latest .
docker build --platform linux/amd64 -t xpdf-serverless:0.3.7 .
Symptoms: TimeoutExpired error after 45-300 seconds
Causes:
Fixes:
aws lambda update-function-configuration --function-name xpdf --memory-size 3008 --region eu-west-1
aws lambda update-function-configuration --function-name xpdf --timeout 900 --region eu-west-1
Cause: LibreOffice trying to write to home directory
Fix: Ensure ENV HOME=/tmp is set in Dockerfile (already configured)
Check S3 event notification:
aws s3api get-bucket-notification-configuration --bucket YOUR-BUCKET-NAME --region eu-west-1
Ensure filters match your file types (.docx, .xlsx, .pptx)
aws logs tail /aws/lambda/xpdf --since 10m --region eu-west-1
s3:GetObject and s3:PutObject permissionsCause: S3 notification triggering on PDF outputs
Fix: Ensure S3 filters exclude .pdf files (see Step 8 configuration with suffix filters)
See docs/dev-guide.md for complete AWS resource cleanup instructions.
S3 Upload → S3 Event → Lambda Invocation → Download File
↓
LibreOffice Convert
↓
Upload PDF to S3
Handler: handler.py
Docker Image: Dockerfile
AI Agent Instructions: .github/copilot-instructions.md
docker run -v $(pwd)/test_documents:/tmp/test_documents libreoffice-service:latest python3 handler.py /tmp/test_documents/file_example_XLSX_10.xlsx
Content type
Image
Digest
sha256:5db3be55b…
Size
751.4 MB
Last updated
8 months ago
docker pull datapsycho/xpdf-serverless