Sign inSign up

otakulabz/docuflow

By otakulabz

•Updated 5 months ago

Image
0

1.1K

otakulabz/docuflow repository overview

⁠doc-parser

A fast, local document parsing service that converts text-based documents into structured Markdown and RAG-ready chunks.

Built to feed vector stores. Exposed as both a REST API and an MCP server so AI agents can call it directly.


⁠What it does

Takes a document (PDF, EPUB, DOCX, TXT, MD, HTML) and returns:

  • Structured Markdown — headings detected from font size, style names, or tag structure depending on format
  • RAG-ready chunks — 500-word chunks with 50-word overlap, section-aware, vector-store ready
  • Stats — page count, word count, sections detected, elapsed time

For PDFs and EPUBs, structure is detected by calibrating font sizes across the document — no hardcoded thresholds. The most common font size becomes "body text" and everything proportionally larger becomes a heading.


⁠Supported formats

ExtensionEngineHow structure is detected
pdf, epub, xpsPyMuPDFFont size calibration across all spans
docxpython-docxWord heading styles (Heading 1, Heading 2, etc.)
md, markdownregex# ## ### prefix matching
txt, textheuristicDouble-newline paragraph splits + ALL-CAPS short lines → H2
html, htmBeautifulSoup<h1>–<h5>, <p>, <li>, <blockquote> tags

⁠Architecture

/home/rag-demo/
├── server/
│   ├── parser.py       ← core parsing module (pure functions, no I/O)
│   ├── api.py          ← FastAPI REST server  (port 8765)
│   └── mcp_server.py   ← MCP SSE server       (port 8766)
├── requirements.txt
├── start.sh            ← starts both servers
├── Dockerfile
└── docker-compose.yml

parser.py is the core — both servers import from it. It has zero I/O and no framework dependencies, so it's easy to test and reuse.

⁠Data flow
document bytes
    │
    ▼
parse_document_bytes(content, filetype, title)
    │
    ├─ _calibrate_fonts()      (PDF/EPUB only — scan all spans, find modal size)
    ├─ _extract_*_structure()  (format-specific section extraction)
    ├─ _sections_to_markdown() (assemble Markdown with # heading levels)
    └─ _sections_to_chunks()   (500-word chunks, 50-word overlap)
    │
    ▼
ParseResult
    ├─ .title     — document title
    ├─ .markdown  — full structured Markdown string
    ├─ .chunks    — list of chunk dicts (id, section, content, markdown, word_count, ...)
    ├─ .sections  — list of detected sections (level, title, page_start)
    └─ .stats     — ParseStats (pages, words, chunks, elapsed_sec, ...)
⁠Chunk format

Each chunk is a dict ready to pass to an embedding API:

{
  "id": "chunk_00042",
  "doc_title": "Own The Day Own Your Life",
  "section": "Morning Hydration",
  "level": 2,
  "content": "raw text content here...",
  "markdown": "## Morning Hydration\n\nraw text content here...",
  "page_start": 34,
  "word_count": 498,
  "char_count": 2941
}

⁠Setup

cd /home/rag-demo
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

⁠Running

bash start.sh
FastAPI REST  →  http://0.0.0.0:8765
MCP SSE       →  http://0.0.0.0:8766/sse
Swagger docs  →  http://0.0.0.0:8765/docs
⁠Individually
# FastAPI only
python3 -m uvicorn server.api:app --host 0.0.0.0 --port 8765

# MCP only
MCP_HOST=0.0.0.0 MCP_PORT=8766 python3 -m server.mcp_server
⁠Docker
docker compose up --build

⁠REST API

MethodEndpointDescription
GET/healthHealth check
GET/formatsList supported formats and engines
POST/parseFile upload → Markdown + chunks
POST/parse/urlURL → Markdown + chunks
POST/extractFile upload → raw text only
POST/extract/urlURL → raw text only

Format is auto-detected from the file extension. Pass format in the request body to override.

Parse a file:

curl -X POST http://localhost:8765/parse \
  -F "[email protected]"

Parse from URL:

curl -X POST http://localhost:8765/parse/url \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/book.epub", "title": "My Book"}'

Get JSONL chunks (MCP tool):

curl -X POST http://localhost:8765/parse/url \
  -H "Content-Type: application/json" \
  -d '{"url": "https://cdn.discord.com/.../report.pdf"}'

Full interactive docs at http://localhost:8765/docs.


⁠MCP server

The MCP server runs on port 8766 with SSE transport. Add it to any MCP-compatible client:

{
  "mcpServers": {
    "doc-parser": {
      "url": "http://<your-server-ip>:8766/sse"
    }
  }
}
⁠Tools
ToolDescription
parse_document_url(url, title?, format?)Parse any supported doc from URL → Markdown + chunks
parse_document_base64(content_base64, filename)Parse from raw base64 bytes → Markdown + chunks
extract_text_url(url, format?)Fast raw text only — no structure, no chunking
get_chunks_jsonl(url, title?, format?)Returns JSONL string — one chunk per line, pipe to vector store
parse_pdf_urlAlias for parse_document_url with format="pdf"
parse_pdf_base64Alias for parse_document_base64
⁠Agent usage example

Once registered, an agent can call:

parse_document_url(
  url="https://cdn.discordapp.com/attachments/.../book.pdf",
  title="Own The Day"
)

And get back the full structured Markdown + all chunks — ready to embed and store.


⁠Performance

Tested on "Own The Day, Own Your Life" (399 pages, 9.7 MB):

Pages:      399
Characters: 739,215
Words:      ~123,000
Time:       0.694s
Speed:      575 pages/sec

PyMuPDF reads the embedded text layer directly — it's not doing OCR. Speed scales linearly with page count. For scanned image PDFs, you'd need a vision model instead.


⁠Using parser.py directly

The parser module is importable independently:

from server.parser import parse_document_bytes, extract_text_bytes

# Parse a PDF
with open("book.pdf", "rb") as f:
    result = parse_document_bytes(f.read(), filetype="pdf", title="My Book")

print(result.stats)          # ParseStats(pages_total=399, chunks=248, ...)
print(result.markdown[:500]) # # My Book\n\n---\n\n## Chapter One...
print(result.chunks[0])      # {id: "chunk_00000", section: "...", ...}

# Fast raw extraction (no structure)
with open("book.pdf", "rb") as f:
    data = extract_text_bytes(f.read(), filetype="pdf")

print(data["words"])       # 123893
print(data["elapsed_sec"]) # 0.312

⁠Environment variables

VariableDefaultDescription
MCP_HOST0.0.0.0MCP server bind host
MCP_PORT8766MCP server port

Docker Compose


##########################
# PDF RAG Stack
##########################

services:
  pdf-api:
    image: otakulabz/docuflow:latest
    pull_policy: always
    container_name: ${COMPOSE_PROJECT_NAME:-pdf}-api
    command: uvicorn server.api:app --host 0.0.0.0 --port ${API_PORT:-8765}
    environment:
      - API_PORT=${API_PORT:-8765}
      - LOG_LEVEL=${LOG_LEVEL:-info}
      - PYTHONPATH=/app
    restart: unless-stopped
    working_dir: /app
    networks:
      - nat
#############################################################################################################
  pdf-mcp:
    image: otakulabz/docuflow:latest
    container_name: ${COMPOSE_PROJECT_NAME:-pdf}-mcp
    pull_policy: always
    command: python -m server.mcp_server
    environment:
      - MCP_PORT=${MCP_PORT:-8766}
      - LOG_LEVEL=${LOG_LEVEL:-info}
      - PYTHONPATH=/app
    restart: unless-stopped
    working_dir: /app
    networks:
      - nat
##########################
# Networks
##########################
networks:
  nat:
    external: true

The FastAPI port is set via the uvicorn command in start.sh or docker-compose.yml.

Tag summary

Content type

Image

Digest

sha256:879559bee…

Size

128.7 MB

Last updated

5 months ago

docker pull otakulabz/docuflow