A fast, local document parsing service that converts text-based documents into structured Markdown and RAG-ready chunks.
Built to feed vector stores. Exposed as both a REST API and an MCP server so AI agents can call it directly.
Takes a document (PDF, EPUB, DOCX, TXT, MD, HTML) and returns:
For PDFs and EPUBs, structure is detected by calibrating font sizes across the document — no hardcoded thresholds. The most common font size becomes "body text" and everything proportionally larger becomes a heading.
| Extension | Engine | How structure is detected |
|---|---|---|
pdf, epub, xps | PyMuPDF | Font size calibration across all spans |
docx | python-docx | Word heading styles (Heading 1, Heading 2, etc.) |
md, markdown | regex | # ## ### prefix matching |
txt, text | heuristic | Double-newline paragraph splits + ALL-CAPS short lines → H2 |
html, htm | BeautifulSoup | <h1>–<h5>, <p>, <li>, <blockquote> tags |
/home/rag-demo/
├── server/
│ ├── parser.py ← core parsing module (pure functions, no I/O)
│ ├── api.py ← FastAPI REST server (port 8765)
│ └── mcp_server.py ← MCP SSE server (port 8766)
├── requirements.txt
├── start.sh ← starts both servers
├── Dockerfile
└── docker-compose.yml
parser.py is the core — both servers import from it. It has zero I/O and no framework dependencies, so it's easy to test and reuse.
document bytes
│
▼
parse_document_bytes(content, filetype, title)
│
├─ _calibrate_fonts() (PDF/EPUB only — scan all spans, find modal size)
├─ _extract_*_structure() (format-specific section extraction)
├─ _sections_to_markdown() (assemble Markdown with # heading levels)
└─ _sections_to_chunks() (500-word chunks, 50-word overlap)
│
▼
ParseResult
├─ .title — document title
├─ .markdown — full structured Markdown string
├─ .chunks — list of chunk dicts (id, section, content, markdown, word_count, ...)
├─ .sections — list of detected sections (level, title, page_start)
└─ .stats — ParseStats (pages, words, chunks, elapsed_sec, ...)
Each chunk is a dict ready to pass to an embedding API:
{
"id": "chunk_00042",
"doc_title": "Own The Day Own Your Life",
"section": "Morning Hydration",
"level": 2,
"content": "raw text content here...",
"markdown": "## Morning Hydration\n\nraw text content here...",
"page_start": 34,
"word_count": 498,
"char_count": 2941
}
cd /home/rag-demo
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
bash start.sh
FastAPI REST → http://0.0.0.0:8765
MCP SSE → http://0.0.0.0:8766/sse
Swagger docs → http://0.0.0.0:8765/docs
# FastAPI only
python3 -m uvicorn server.api:app --host 0.0.0.0 --port 8765
# MCP only
MCP_HOST=0.0.0.0 MCP_PORT=8766 python3 -m server.mcp_server
docker compose up --build
| Method | Endpoint | Description |
|---|---|---|
GET | /health | Health check |
GET | /formats | List supported formats and engines |
POST | /parse | File upload → Markdown + chunks |
POST | /parse/url | URL → Markdown + chunks |
POST | /extract | File upload → raw text only |
POST | /extract/url | URL → raw text only |
Format is auto-detected from the file extension. Pass format in the request body to override.
Parse a file:
curl -X POST http://localhost:8765/parse \
-F "[email protected]"
Parse from URL:
curl -X POST http://localhost:8765/parse/url \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/book.epub", "title": "My Book"}'
Get JSONL chunks (MCP tool):
curl -X POST http://localhost:8765/parse/url \
-H "Content-Type: application/json" \
-d '{"url": "https://cdn.discord.com/.../report.pdf"}'
Full interactive docs at http://localhost:8765/docs.
The MCP server runs on port 8766 with SSE transport. Add it to any MCP-compatible client:
{
"mcpServers": {
"doc-parser": {
"url": "http://<your-server-ip>:8766/sse"
}
}
}
| Tool | Description |
|---|---|
parse_document_url(url, title?, format?) | Parse any supported doc from URL → Markdown + chunks |
parse_document_base64(content_base64, filename) | Parse from raw base64 bytes → Markdown + chunks |
extract_text_url(url, format?) | Fast raw text only — no structure, no chunking |
get_chunks_jsonl(url, title?, format?) | Returns JSONL string — one chunk per line, pipe to vector store |
parse_pdf_url | Alias for parse_document_url with format="pdf" |
parse_pdf_base64 | Alias for parse_document_base64 |
Once registered, an agent can call:
parse_document_url(
url="https://cdn.discordapp.com/attachments/.../book.pdf",
title="Own The Day"
)
And get back the full structured Markdown + all chunks — ready to embed and store.
Tested on "Own The Day, Own Your Life" (399 pages, 9.7 MB):
Pages: 399
Characters: 739,215
Words: ~123,000
Time: 0.694s
Speed: 575 pages/sec
PyMuPDF reads the embedded text layer directly — it's not doing OCR. Speed scales linearly with page count. For scanned image PDFs, you'd need a vision model instead.
The parser module is importable independently:
from server.parser import parse_document_bytes, extract_text_bytes
# Parse a PDF
with open("book.pdf", "rb") as f:
result = parse_document_bytes(f.read(), filetype="pdf", title="My Book")
print(result.stats) # ParseStats(pages_total=399, chunks=248, ...)
print(result.markdown[:500]) # # My Book\n\n---\n\n## Chapter One...
print(result.chunks[0]) # {id: "chunk_00000", section: "...", ...}
# Fast raw extraction (no structure)
with open("book.pdf", "rb") as f:
data = extract_text_bytes(f.read(), filetype="pdf")
print(data["words"]) # 123893
print(data["elapsed_sec"]) # 0.312
| Variable | Default | Description |
|---|---|---|
MCP_HOST | 0.0.0.0 | MCP server bind host |
MCP_PORT | 8766 | MCP server port |
Docker Compose
##########################
# PDF RAG Stack
##########################
services:
pdf-api:
image: otakulabz/docuflow:latest
pull_policy: always
container_name: ${COMPOSE_PROJECT_NAME:-pdf}-api
command: uvicorn server.api:app --host 0.0.0.0 --port ${API_PORT:-8765}
environment:
- API_PORT=${API_PORT:-8765}
- LOG_LEVEL=${LOG_LEVEL:-info}
- PYTHONPATH=/app
restart: unless-stopped
working_dir: /app
networks:
- nat
#############################################################################################################
pdf-mcp:
image: otakulabz/docuflow:latest
container_name: ${COMPOSE_PROJECT_NAME:-pdf}-mcp
pull_policy: always
command: python -m server.mcp_server
environment:
- MCP_PORT=${MCP_PORT:-8766}
- LOG_LEVEL=${LOG_LEVEL:-info}
- PYTHONPATH=/app
restart: unless-stopped
working_dir: /app
networks:
- nat
##########################
# Networks
##########################
networks:
nat:
external: true
The FastAPI port is set via the uvicorn command in start.sh or docker-compose.yml.
Content type
Image
Digest
sha256:879559bee…
Size
128.7 MB
Last updated
5 months ago
docker pull otakulabz/docuflow