Sign inSign up

jparkerweb/semantic-chunking

By jparkerweb

•Updated about 2 months ago

Self-hosted semantic text chunking for RAG/LLMs — REST API + Web UI, local ONNX embeddings

Image
Machine learning & AI
0

1.5K

jparkerweb/semantic-chunking repository overview

⁠semantic-chunking

Self-hosted semantic text chunking for RAG and LLM pipelines. Splits large documents into meaningful, context-preserving chunks using sentence similarity and local ONNX embeddings — no external API required. Ships a REST API and an optional Web UI.

Powered by the semantic-chunking⁠ npm library (internally backed by embedding-utils⁠).

⁠Features

  • 🧩 Three chunking modes: chunkit (semantic), cramit (dense packing), sentenceit (sentence split)
  • 🔒 Runs fully locally — on-device ONNX embeddings, models cached to a volume
  • 🌐 REST API (port 3001) + optional Web UI (port 3000)
  • 🔑 Optional Bearer-token auth on all /api/* routes
  • ❤️ Built-in health check at /api/health

⁠Quick Start

API server (default):

docker run -d --name semantic-chunking-api \
  -p 3001:3001 \
  -v "$(pwd)/models:/app/models" \
  jparkerweb/semantic-chunking:latest

Web UI (override the default command):

docker run -d --name semantic-chunking-webui \
  -p 3000:3000 -e PORT=3000 \
  -v "$(pwd)/models:/app/models" \
  jparkerweb/semantic-chunking:latest node webui/server.js

Then open the Web UI at http://localhost:3000 or call the API at http://localhost:3001.

⁠docker-compose

services:
  api:
    image: jparkerweb/semantic-chunking:latest
    ports:
      - "3001:3001"
    volumes:
      - ./models:/app/models
    environment:
      - PORT=3001
      # - API_AUTH_TOKEN=your-secret-token-here
    restart: unless-stopped

⁠Example request

curl -X POST http://localhost:3001/api/chunkit \
  -H "Content-Type: application/json" \
  -d '{"documents":[{"document_name":"doc1","document_text":"Your long text here..."}]}'

⁠API endpoints

MethodPathDescription
GET/api/healthHealth check
GET/api/versionLibrary/API version
POST/api/chunkitSemantic, similarity-based chunking
POST/api/cramitDense token-packed chunking
POST/api/sentenceitSentence-level splitting

⁠Ports

PortService
3001REST API (default)
3000Web UI (optional)

⁠Volumes

PathPurpose
/app/modelsPersist downloaded ONNX models between restarts

Mount ./models:/app/models so the embedding model is downloaded once and reused.

⁠Environment variables

VariableDefaultDescription
PORT3001Port the server listens on
NODE_ENVproductionNode environment
API_AUTH_TOKEN(unset)If set, all /api/* routes require Authorization: Bearer <token>

⁠Tags

  • latest — most recent release

Tag summary

Content type

Image

Digest

sha256:0f7b757b2…

Size

717.3 MB

Last updated

about 2 months ago

docker pull jparkerweb/semantic-chunking