Sign inSign up

jchristn77/documentatom-mcp

By jchristn77

•Updated 4 days ago

API server for breaking docs into atoms for text processing, analysis, and AI.

Image
Machine learning & AI
Data science
0

3.1K

jchristn77/documentatom-mcp repository overview

⁠DocumentAtom

DocumentAtom provides a light, fast library for breaking input documents into constituent parts (atoms), useful for text processing, analysis, and artificial intelligence.

DocumentAtom requires that Tesseract v5.0 be installed on the host. This is required as certain document types can have embedded images which are parsed using OCR via Tesseract.

PackageVersionDownloads
DocumentAtom.CsvNuGet VersionNuGet
DocumentAtom.ExcelNuGet VersionNuGet
DocumentAtom.HtmlNuGet VersionNuGet
DocumentAtom.ImageNuGet VersionNuGet
DocumentAtom.JsonNuGet VersionNuGet
DocumentAtom.MarkdownNuGet VersionNuGet
DocumentAtom.PdfNuGet VersionNuGet
DocumentAtom.PowerPointNuGet VersionNuGet
DocumentAtom.OcrNuGet VersionNuGet
DocumentAtom.RichTextNuGet VersionNuGet
DocumentAtom.TextNuGet VersionNuGet
DocumentAtom.TypeDetectionNuGet VersionNuGet
DocumentAtom.WordNuGet VersionNuGet
DocumentAtom.XmlNuGet VersionNuGet

⁠New in v1.1.x

  • Hierarchical atomization (see BuildHierarchy in settings) - heading-based for markdown/HTML/Word, page-based for PowerPoint
  • Support for CSV, JSON, and XML documents
  • MCP server (DocumentAtom.McpServer) for exposing DocumentAtom operations via Model Context Protocol to AI assistants
  • Dependency updates and fixes

⁠Motivation

Parsing documents and extracting constituent parts is one part science and one part black magic. If you find ways to improve processing and extraction in any way that is horizontally useful, I'd would love your feedback on ways to make this library more accurate, more useful, faster, and overall better. My goal in building this library is to make it easier to analyze input data assets and make them more consumable by other systems including analytics and artificial intelligence.

⁠Bugs, Quality, Feedback, or Enhancement Requests

Please feel free to file issues, enhancement requests, or start discussions about use of the library, improvements, or fixes.

⁠Types Supported

DocumentAtom supports the following input file types:

  • CSV
  • HTML
  • JSON
  • Markdown
  • Microsoft Word (.docx)
  • Microsoft Excel (.xlsx)
  • Microsoft PowerPoint (.pptx)
  • PNG images (requires Tesseract on the host)
  • PDF
  • Rich text (.rtf)
  • Text
  • XML

⁠Simple Example

Refer to the various Test projects for working examples.

The following example shows processing a markdown (.md) file.

using DocumentAtom.Core.Atoms;
using DocumentAtom.Markdown;

MarkdownProcessorSettings settings = new MarkdownProcessorSettings();
MarkdownProcessor processor = new MarkdownProcessor(_Settings);
foreach (Atom atom in processor.Extract(filename))
    Console.WriteLine(atom.ToString());

⁠Atom Types

DocumentAtom parses input data assets into a variety of Atom objects. Each Atom includes top-level metadata including:

  • ParentGUID - globally-unique identifier of the parent atom, or, null
  • GUID - globally-unique identifier
  • Type - including Text, Image, Binary, Table, and List
  • PageNumber - where available; some document types do not explicitly indicate page numbers, and page numbers are inferred when rendered
  • Position - the ordinal position of the Atom, relative to others
  • Length - the length of the Atom's content
  • MD5Hash - the MD5 hash of the Atom content
  • SHA1Hash - the SHA1 hash of the Atom content
  • SHA256Hash - the SHA256 hash of the Atom content
  • Quarks - sub-atomic particles created from the Atom content, for instance, when chunking text

The AtomBase class provides the aforementioned metadata, and several type-specific Atoms are returned from the various processors, including:

  • BinaryAtom - includes a Bytes property
  • DocxAtom - includes Text, HeaderLevel, UnorderedList, OrderedList, Table, and Binary properties
  • ImageAtom - includes BoundingBox, Text, UnorderedList, OrderedList, Table, and Binary properties
  • MarkdownAtom - includes Formatting, Text, UnorderedList, OrderedList, and Table properties
  • PdfAtom - includes BoundingBox, Text, UnorderedList, OrderedList, Table, and Binary properties
  • PptxAtom - includes Title, Subtitle, Text, UnorderedList, OrderedList, Table, and Binary properties
  • TableAtom - includes Rows, Columns, Irregular, and Table properties
  • TextAtom - includes Text
  • XlsxAtom - includes SheetName, CellIdentifier, Text, Table, and Binary properties

Table objects inside of Atom objects are always presented as SerializableDataTable objects (see SerializableDataTable⁠ for more information) to provide simple serialization and conversion to native System.Data.DataTable objects.

⁠Underlying Libraries

DocumentAtom is built on the shoulders of several libraries, without which, this work would not be possible.

Each of these libraries were integrated as NuGet packages, and no source was included or modified from these packages.

My libraries used within DocumentAtom:

⁠RESTful API and Docker

Run the DocumentAtom.Server project to start a RESTful server listening on localhost:8000. Modify the documentatom.json file to change the webserver, logging, or Tesseract settings. Alternatively, you can pull jchristn/documentatom from Docker Hub⁠. Refer to the Docker directory in the project for assets for running in Docker.

Refer to the Postman collection for examples exercising the APIs.

⁠Running Locally
cd src/DocumentAtom.Server
dotnet run
⁠Running with Docker
  1. Pull the image from Docker Hub:
docker pull jchristn/documentatom:v1.1.0
  1. Create a documentatom.json configuration file (see Docker/documentatom.json for an example)

  2. Run the container:

# Windows
docker run -p 8000:8000 -v .\documentatom.json:/app/documentatom.json -v .\logs\:/app/logs/ jchristn/documentatom:v1.1.0

# Linux/macOS
docker run -p 8000:8000 -v ./documentatom.json:/app/documentatom.json -v ./logs/:/app/logs/ jchristn/documentatom:v1.1.0

Alternatively, use the provided scripts in the Docker directory:

# Windows
Dockerrun.bat v1.1.0

# Linux/macOS
IMG_TAG=v1.1.0 ./Dockerrun.sh

⁠MCP Server and Docker

The DocumentAtom.McpServer project provides a Model Context Protocol (MCP)⁠ server that exposes DocumentAtom operations to AI assistants and LLM-based tools. The MCP server acts as a front-end to the DocumentAtom.Server RESTful API, enabling AI agents to process documents via standardized MCP tool calls.

The MCP server supports three transport protocols:

  • HTTP: JSON-RPC over HTTP at /rpc (default port 8200)
  • TCP: Raw TCP socket connection (default port 8201)
  • WebSocket: WebSocket connection at /mcp (default port 8202)
⁠Prerequisites

The MCP server requires a running DocumentAtom.Server instance. Configure the endpoint in documentatom.json:

{
  "DocumentAtom": {
    "Endpoint": "http://localhost:8000",
    "AccessKey": null
  }
}
⁠Running Locally
cd src/DocumentAtom.McpServer
dotnet run

Command-line options:

  • --config=<file> - Specify settings file path (default: ./documentatom.json)
  • --showconfig - Display configuration and exit
  • --help, -h - Show help message
⁠Running with Docker
  1. Pull the image from Docker Hub:
docker pull jchristn/documentatom-mcp:v1.1.0
  1. Create a documentatom.json configuration file with MCP server settings:
{
  "Logging": {
    "LogDirectory": "./logs/",
    "LogFilename": "documentatom-mcp.log",
    "ConsoleLogging": true,
    "EnableColors": true,
    "MinimumSeverity": 0
  },
  "DocumentAtom": {
    "Endpoint": "http://host.docker.internal:8000",
    "AccessKey": null
  },
  "Http": {
    "Hostname": "0.0.0.0",
    "Port": 8200
  },
  "Tcp": {
    "Address": "0.0.0.0",
    "Port": 8201
  },
  "WebSocket": {
    "Hostname": "0.0.0.0",
    "Port": 8202
  },
  "Storage": {
    "BackupsDirectory": "./backups/",
    "TempDirectory": "./temp/"
  }
}
  1. Run the container:
# Windows
docker run -p 8200:8200 -p 8201:8201 -p 8202:8202 ^
  -v .\documentatom.json:/app/documentatom.json ^
  -v .\logs\:/app/logs/ ^
  -v .\temp\:/app/temp/ ^
  -v .\backups\:/app/backups/ ^
  jchristn/documentatom-mcp:v1.1.0

# Linux/macOS
docker run -p 8200:8200 -p 8201:8201 -p 8202:8202 \
  -v ./documentatom.json:/app/documentatom.json \
  -v ./logs/:/app/logs/ \
  -v ./temp/:/app/temp/ \
  -v ./backups/:/app/backups/ \
  jchristn/documentatom-mcp:v1.1.0

Alternatively, use the provided scripts in src/DocumentAtom.McpServer:

# Windows
Dockerrun.bat v1.0.0

# Linux/macOS
IMG_TAG=v1.0.0 ./Dockerrun.sh
⁠Environment Variables

The MCP server supports the following environment variables to override configuration:

VariableDescription
DOCUMENTATOM_ENDPOINTDocumentAtom server endpoint URL
DOCUMENTATOM_ACCESS_KEYAccess key for authentication
MCP_HTTP_HOSTNAMEHTTP server hostname
MCP_HTTP_PORTHTTP server port
MCP_TCP_ADDRESSTCP server address
MCP_TCP_PORTTCP server port
MCP_WEBSOCKET_HOSTNAMEWebSocket server hostname
MCP_WEBSOCKET_PORTWebSocket server port
CONSOLE_LOGGINGEnable console logging (1 or 0)
⁠Building Docker Images

To build the Docker images locally:

# Build DocumentAtom.Server image
cd Docker
Dockerbuild.bat v1.1.0 0  # 0 = don't push, 1 = push to Docker Hub

# Build DocumentAtom.McpServer image (from src directory)
cd src
docker buildx build -f DocumentAtom.McpServer/Dockerfile --platform linux/amd64,linux/arm64/v8 --tag jchristn/documentatom-mcp:v1.1.0 --push .

⁠Version History

Please refer to CHANGELOG.md for version history.

⁠Thanks

Special thanks to iconduck.com and the content authors for producing this icon⁠.

Tag summary

Content type

Image

Digest

sha256:39c4758f0…

Size

341 MB

Last updated

4 days ago

docker pull jchristn77/documentatom-mcp