Sign inSign up

molandev/mineru-converter-api

By molandev

•Updated 7 months ago

Mineru HTTP API (port 10996) for PDF conversion to Markdown (CPU-only mode)

Image
Machine learning & AI
Developer tools
0

587

molandev/mineru-converter-api repository overview

⁠MinerU PDF Converter Service - Overview

ā šŸ“¦ Source Code

ā āš ļø Important Notice

This image runs in CPU-only mode with pipeline backend.

For GPU acceleration, please use the official MinerU image⁠.

  • CPU Mode: Suitable for moderate workloads, no GPU required
  • Pipeline Backend: Optimized for general document processing
  • GPU Mode: Visit official repository for high-performance GPU support

This image only supports PDF files.

  • Does NOT support .docx or .doc Word documents
  • Does NOT support .pptx, .xlsx, or other Office formats

For other document types, use our other converters:

Document TypeRecommended Image
DOCX, PPTX, XLSX → Markdownmolandev/markitdown-converter-api:latest
DOC → DOCX/PDFmolandev/libreoffice-converter-api:latest

ā šŸš€ What Is It?

A high-precision PDF parsing service powered by MinerU with advanced OCR capabilities. Built on GraalVM Native Image for cloud-native deployment, it transforms complex PDF documents into structured Markdown or JSON format with exceptional accuracy.

⁠✨ Core Features

ā šŸ”„ Superior PDF Parsing
  • OCR Engine: Extract text from scanned documents and images
  • Layout Preservation: Maintains document structure and formatting
  • Table Detection: Intelligent table extraction with proper formatting
  • Image Extraction: Preserves images with organized output
ā šŸ›”ļø Production-Ready Reliability
  • Smart Concurrency Control: Semaphore-based rate limiting prevents system overload
  • Auto-Cleanup Mechanism: Scheduled task automatically purges temporary files
  • Graceful Degradation: Queuing mechanism with configurable timeouts
ā šŸŽÆ Flexible Output Formats
  • Markdown: Clean, LLM-friendly text output
  • JSON: Structured data with full metadata
  • ZIP Archive: All output files in one package (md, json, pdf variants, images)

ā šŸ“¦ Output Files (ZIP Format)

FileDescription
result.mdExtracted Markdown content
result_content_list.jsonContent structure list
result_middle.jsonIntermediate processing data
result_model.jsonModel output data
result_layout.pdfLayout analysis PDF
result_origin.pdfOriginal PDF copy
result_span.pdfSpan annotation PDF
images/Extracted images folder

ā šŸŽ® Quick Start

⁠Pull from Docker Hub
docker pull molandev/mineru-converter-api:latest
⁠Run in Seconds
docker run -d \
  --name mineru-converter \
  -p 10996:10996 \
  -v /data/temp:/app/temp \
  -e CONVERTER_MAX_CONCURRENT=5 \
  -e CONVERTER_RETENTION_MINUTES=30 \
  molandev/mineru-converter-api:latest
ā šŸ“š API Documentation

Access the complete API guide at: http://localhost:10996/help

⁠Convert Your First PDF
# Get Markdown (default)
curl -F "[email protected]" \
     http://localhost:10996/convert/mineru/upload \
     -o output.md

# Get JSON
curl -F "[email protected]" \
     "http://localhost:10996/convert/mineru/upload?outputFormat=json" \
     -o output.json

# Get all files (ZIP)
curl -F "[email protected]" \
     "http://localhost:10996/convert/mineru/upload?outputFormat=zip" \
     -o output.zip

# Or provide a URL
curl -X POST "http://localhost:10996/convert/mineru/url?fileUrl=https://example.com/doc.pdf&outputFormat=zip" \
     -o output.zip

ā āš™ļø Configuration

ParameterDefaultDescription
CONVERTER_MAX_CONCURRENT5Max parallel conversions (MinerU is resource-intensive)
CONVERTER_RETENTION_MINUTES30Temp file retention
CONVERTER_SCHEDULE_MINUTES5Cleanup interval
SERVER_PORT10996Service port

ā šŸ”§ Advanced Usage

⁠Resource-Constrained Environment
docker run -e CONVERTER_MAX_CONCURRENT=2 \
           -e CONVERTER_RETENTION_MINUTES=60 \
           molandev/mineru-converter-api:latest
⁠High-Performance Setup
docker run -e CONVERTER_MAX_CONCURRENT=10 \
           --memory=8g \
           --cpus=4 \
           molandev/mineru-converter-api:latest

ā šŸ’” Why Choose This?

āœ… OCR-Powered - Extract text from scanned PDFs with high accuracy
āœ… Multiple Output Formats - Markdown, JSON, or complete ZIP archive
āœ… Resource-Efficient - GraalVM Native Image for minimal footprint
āœ… Production-Ready - Smart concurrency control & auto-cleanup
āœ… Cloud-Native - Docker-first design, Kubernetes-ready

ā šŸŽÆ Perfect For

  • AI Knowledge Base Pipelines: Convert PDFs to LLM-friendly Markdown
  • Document Digitization: OCR for scanned documents and archives
  • Data Extraction: Structured JSON output for downstream processing
  • Research & Analysis: Extract tables, figures, and structured content

ā šŸ“Š API Reference

⁠Upload Endpoint
POST /convert/mineru/upload
ParameterTypeRequiredDescription
fileFileYesPDF file to convert
outputFormatStringNoOutput format: md (default), json, zip
⁠URL Endpoint
POST /convert/mineru/url
ParameterTypeRequiredDescription
fileUrlStringYesRemote PDF URL
outputFormatStringNoOutput format: md (default), json, zip

Built with Spring Boot 3.5.6 + GraalVM 25 + MinerU

Precision PDF parsing with OCR intelligence. šŸ“Š

Tag summary

Content type

Image

Digest

sha256:d266fe868…

Size

7.1 GB

Last updated

7 months ago

docker pull molandev/mineru-converter-api