A high-performancegraph-based RAG
1.3K
High-performance Rust + Python engine for graph-based RAG, combining vector search, graph traversal, and LLM reasoning in a single system.
Query → Embedding → HNSW → Graph (BFS) → Ranking → LLM
Engine + text plugin in one container, communicating over a Unix socket. Bundles text embedding/extraction/generation only — no image/CLIP support (use the multi-container HTTP setup below for that).
docker run -p 8000:8000 \
-v $(pwd)/data:/data \
--env-file .env \
khapu2906/linkingmem:latest
Bring your own embedding / extraction / generation service (any language, HTTP or Unix socket — see Plugin Interface):
docker run -p 8000:8000 \
-v $(pwd)/data:/data \
--env-file .env \
khapu2906/linkingmem:latest-engine
For text and image queries (visual similarity search, /query/image),
run the engine + text plugin + image plugin as three containers via
docker compose instead of a single image — see
docker-compose.yml
in the source repo.
| Tag | Description |
|---|---|
latest | Alias for the newest vX.Y.Z-full release |
latest-engine | Alias for the newest vX.Y.Z-engine release |
v0.3.0-full | Rust engine + Python text plugin (all-in-one, Unix socket) |
v0.3.0-engine | Rust engine only (bring your own plugin) |
Minimum:
OPENAI_API_KEY=your_api_key
Works with any OpenAI-compatible endpoint — OpenAI, Ollama, Gemini
(compat mode), Groq, LM Studio, vLLM, etc. — by also setting
OPENAI_BASE_URL (defaults to https://api.openai.com/v1).
See full config: 👉 https://github.com/khapu2906/LinkingMem/blob/main/.env.example
curl -X POST http://localhost:8000/query/text \
-H "Content-Type: application/json" \
-d '{"query": "Who works at Acme Corp?"}'
Full API reference: 👉 https://github.com/khapu2906/LinkingMem/blob/main/docs/API_REFERENCE.md
OPENAI_BASE_URL at any compatible provider, no vendor lock-in👉 https://github.com/khapu2906/LinkingMem
The Rust core itself contributes well under 1ms to query latency — graph traversal and vector search are both sub-millisecond even at 100k+ nodes (see full benchmarks).
End-to-end query latency is dominated by your LLM provider's response
time, not the engine — typically 350–900ms with a fast, nearby LLM
endpoint, but this varies significantly (multi-second) with slower or
more distant providers. Run /query/vector or /query/node (no LLM
call) to measure pure retrieval latency in your environment.
Content type
Image
Digest
sha256:476c10b20…
Size
272.8 MB
Last updated
3 months ago
docker pull khapu2906/linkingmem