Sign inSign up

lyellr88/marm-mcp-server

By lyellr88

•Updated about 22 hours ago

Universal MCP Server: intelligent memory, semantic search, & AI agent collaboration support.

Image
Machine learning & AI
1

50K+

lyellr88/marm-mcp-server repository overview

marm-memory - persistent local memory server for AI agents (Model Context Protocol).

⁠marm-memory v2.48.1 - Give your AI Agents a permanent memory in 60 seconds

License Python FastAPI Docker Pulls PyPI Downloads PyPI Version MCP Registry

Discord Publish CodeQL marm-memory MCP server

⁠Quick Start

  1. Install and initialize with your preferred agent profiles:
pip install marm-mcp-server
marm-memory init --g-claude --g-codex --g-gemini

Also available: --g-qwen and --g-kiro. Run without flags to install into your current project folder instead of home

  1. Hand off to your AI companion. Tell your agent:

"Use the marm-init skill to set up MARM."

  1. Interact: Your agent will handle the entire setup (Python/Docker, HTTP/STDIO, keys, and client configs) interactively right inside your chat.

Manual setup

Prefer to wire it up yourself:

Replace "agent" with your client’s CLI command (for example, claude, gemini, or qwen). For Codex, use codex mcp add marm-memory --url http://localhost:8001/mcp⁠ instead.

If you are...Start the serverConnect your MCP client
Solo developer / researchermarm-memory start"agent" mcp add --transport http marm-memory http://localhost:8001/mcp
Private local STDIO usermarm-mcp-stdio"agent" mcp add --transport stdio marm-memory-stdio marm-mcp-stdio
Multiple agents sharing memorymarm-memory start --profile swarm"agent" mcp add --transport http marm-memory http://localhost:8001/mcp
Private high-throughput swarmmarm-memory start --profile swarm-max"agent" mcp add --transport http marm-memory http://localhost:8001/mcp
Trusted private lab/servermarm-memory start --profile trusted"agent" mcp add --transport http marm-memory http://localhost:8001/mcp
  • ⚡ Fastest HTTP Startup: Run marm-memory fast-start-http to spin up the local runtime, launch the console, and open it in your browser immediately.
  • 🖥️ Web Console: Run marm-memory console to view the local UI app instantly (no Node.js required).
  • ⚙️ Lifecycle Management: Manage the background daemon using status, logs --follow, restart, and stop.
  • 💡 Quick Flags: Use --no-console or --no-browser to restrict startups. Run marm-memory --help for full command lists.

⁠Why MARM Memory

Your AI forgets everything. MARM Memory doesn't.

marm-memory gives your agents a private, shared memory for the context that normally gets lost between chats: decisions, research, fixes, notes, and project history. Switch from Claude Code to Codex or Gemini without losing the context already gathered.

It brings three things together:

  • 🧠 Core Memory (7 tools) stores conversations, notes, notebook entries, and summaries so they stay searchable.
  • 💻 Code Graph (5 tools) maps your repository so agents can find symbols, follow code paths, and understand the project without rereading it all. Point it at a repo once, and it keeps itself current as you work.
  • 🧩 Concept Graph (2 tools) connects people, decisions, errors, and ideas from your stored memories, with links back to relevant code when available. It builds itself as you store memories.
⁠How It Works
LayerWhat it doesWhy it matters
Memory modelSessions, structured logs, notebooks, summaries, and semantic memoriesKeeps project history searchable instead of trapped in one chat
Scale layerSQLite WAL mode, connection pooling, serialized write queue, and HTTP rate-limit presetsLets one server support solo use, multi-agent work, and swarm-style bursts
Intelligence layerFTS filter, semantic re-rank, bounded semantic fallback, auto-classification, write-time consolidation, and compaction candidatesKeeps recall useful as memory grows instead of letting duplicates pile up
Code graph layerRepo indexing, symbol lookup, call tracing, architecture overview, and change-impact analysisGives agents project structure without rereading the whole codebase
Concept graph layerEntity and relationship extraction from stored memories, with links back into the code graphConnects decisions, errors, tools, and people across sessions instead of leaving them as flat text
Token layerLightweight 7-tool core surface (14 total with bundled graph tools), semantic re-rank before retrieval, and write-time deduplicationReduces tokens sent to the model on every recall and cost stays predictable as memory scales
Deployment layerPip, Docker, STDIO, HTTP, and managed swarm, swarm-max, and trusted profilesLets you run private local memory or shared multi-agent memory with the same MCP surface
⁠Runtime CLI Commands
marm-memory fast-start-http                # start HTTP, Console, and open the browser
marm-memory start                          # start or reuse the managed HTTP runtime
marm-memory start --profile swarm          # shared multi-agent preset
marm-memory stop                           # stop the managed runtime safely
marm-memory restart                        # restart the managed runtime
marm-memory status                         # inspect runtime, database, queue, and graph status
marm-memory logs --follow                  # follow bounded runtime logs
marm-memory console                        # start or reuse the bundled local Console

Transports and setup

marm-memory http                           # run HTTP in the foreground
marm-memory stdio                          # run the strict local MCP STDIO transport
marm-memory init                           # install the MARM skill into detected agents (project scan)
marm-memory init --g-claude                # install the skill into the home-folder claude directory
marm-memory doctor                         # diagnose the local install
marm-memory key init                       # create or reuse ~/.marm/.env without displaying the key
marm-memory key path                       # print the managed key-file path
marm-memory key reveal                     # explicitly display the managed key
marm-memory console --import-key           # open an authenticated local Console session
marm-memory upgrade --check                # compare the installed package with PyPI
marm-memory uninstall                      # preview package removal; always preserves ~/.marm

Knowledge, projects, and maintenance

marm-memory knowledge status               # Indexers, models, and how far behind automatic indexing is
marm-memory knowledge build --all          # Rebuild the whole concept graph (new memories index themselves)
marm-memory knowledge auto off             # Stop indexing memories automatically (on, off, status)
marm-memory projects list                  # List all tracked workspaces
marm-memory projects index <path>          # Add a repo to the code graph (kept current after that)
marm-memory projects status                # Inspect target repo graph readiness
marm-memory projects auto off              # Stop re-indexing repos automatically (on, off, status)
marm-memory maintenance status             # Check internal database optimization state
marm-memory maintenance embeddings migrate # Upgrade old 384-dim vectors to 512-dim
marm-memory maintenance chunks rechunk     # Recalibrate long memory text splits

Docker commands are documented separately below because they require explicit data mounts, network exposure, and key-handling choices.

⁠Performance & Scaling Benchmarks

MARM is tuned for fast recall first, even as memory grows and long memories are chunked behind the scenes.

These measurements use the fastembed-backed jinaai/jina-embeddings-v2-small-en encoder and a throwaway local SQLite database. Every timed path calls the shipped MARMMemory code, not a benchmark-local reimplementation. Sections 1-4 are timings from a single run of scripts/benchmarking/performance/bench_hotpath.py⁠ on local hardware; absolute milliseconds vary by machine, so treat the scaling shape as the signal. Section 5 is a separate accuracy benchmark (run_eval.py⁠) measuring retrieval rather than speed, and its latest row is a controlled before-and-after, explained there.

⁠1. Retrieval Latency Scaling

End-to-end recall_similar latency (includes query encoding).

Session Size ($N$)Min LatencyMedian Latencyp95 Latency
N = 1007.4 ms7.9 ms9.4 ms
N = 25011.9 ms13.5 ms15.4 ms
N = 50010.9 ms11.8 ms13.4 ms
N = 1,00013.3 ms13.5 ms15.6 ms
N = 2,00017.5 ms18.2 ms19.6 ms
N = 4,00023.8 ms25.9 ms30.9 ms
⁠2. Encoder + Concurrency
  • Cold model load: 893ms
  • Warm encode: median 3.8ms, p95 4.3ms
  • Concurrent recall: 10 gathered recalls completed in 151.5ms vs 176.0ms serial (gather/serial = 0.86). Do not read that as parallelism: repeated runs of this same benchmark land anywhere from 0.63 to 0.86, so the ratio is not stable enough to claim a speedup. The path is serialized around shared encoder and SQLite work by design, and any apparent gain is measurement noise.
⁠3. Write-Time Ingestion Cost
  • Consolidation off: median 6.5ms, p95 7.6ms
  • Consolidation on: median 58.1ms, p95 106.5ms
  • Tradeoff: write-time dedupe/clustering adds 9.0x median cost so recall stays fast, and the store stays cleaner over time. Consolidation is off by default.
⁠4. Recall Scaling: Full Scan vs Production Hybrid

Why recall stays flat as memory grows: Instead of scanning every vector, production recall uses an FTS keyword pre-filter to narrow the candidate pool, then re-ranks using a blended semantic + BM25 + temporal score. Both benchmark columns represent authentic asynchronous code paths timed with precomputed vectors to isolate retrieval speed from raw encoding overhead. Tests alternate execution to ensure completely unbiased cache conditions.

Session Size ($N$)Full Semantic ScanProduction HybridSpeedupFTS candidates
N = 1003.3 ms6.6 ms0.5x85 / 200
N = 50016.3 ms11.6 ms1.4x200 / 200
N = 1,00031.1 ms14.7 ms2.1x200 / 200
N = 2,00063.5 ms19.0 ms3.3x200 / 200
N = 4,000127.2 ms29.1 ms4.4x200 / 200
N = 10,000316.7 ms53.8 ms5.9x200 / 200
⁠5. LoCoMo Retrieval Accuracy

All 10 LoCoMo conversations are ingested through marm_log_entry (5,882 memories), then top-5 marm_smart_recall results are scored against 1,977 evidence-annotated questions. No answer-generation model or LLM judge is involved, so this measures whether the right memory is retrieved, not whether an agent answers correctly with it.

ConfigurationAny evidence hitAll evidence hitMean evidence recall
MiniLM baseline37.5%29.5%not published
Jina v2 Small (v2.29.0)53.0%43.4%47.6%
v2.33.1 through v2.44.362.9 - 63.5%53.1 - 53.5%57.4 - 57.9%
v2.44.4 (log lane fix)69.1 - 69.6%58.2 - 58.6%63.0 - 63.5%
⁠6. vs Competitors: Architecture

MARM targets a specific niche: local-first memory for MCP-connected coding agents, not general personalization memory or a full agent runtime. Here's how it differs architecturally from established names in AI agent memory:

MARMMem0Letta (MemGPT)Zep / Graphitiagentmemory
TypeMemory engine, MCP-nativeMemory layer APIFull agent runtimeTemporal knowledge graphMemory engine, MCP-native
Required infrastructureNo separate data service (embedded SQLite)Vector DB (Qdrant/pgvector)Postgres + vector DBNeo4jSeparate iii-engine runtime
DeploymentLocal-first by default; Docker for shared/remoteCloud API or self-hostedSelf-hosted or cloudCloud or self-hostedLocal-first
Retrieval modelHybrid: FTS5 BM25 exact lane + semantic rerankVector + graph + key-valueVector archival store + agent-managed core memoryTemporal knowledge graph (fact validity windows)BM25 + vector + graph (RRF fusion)
Write captureExplicit tool calls from the connected agentExplicit add() calls (some integrations auto-extract)Agent self-edits its own memoryExplicit API callsHook-based, automatic (no explicit calls needed)
Code structure awarenessBundled code graph + concept graph, fused with memoryNot built inNot built inNot built inNot built in (pairs with a separate project)
Framework lock-inNone (any MCP client)NoneHigh (must run within Letta)NoneNone (any MCP client)

Disclaimers & Accuracy: Competitor landscapes evolve rapidly. The matrix above reflects core architectural traits as of Q3 2026, based on public documentation and READMEs, not internal testing of each system. If any data point regarding an alternative framework has changed or is misrepresented, please open an issue or submit a Pull Request to update the table. We actively welcome corrections from peer maintainers.

⁠MCP Client Setup for HTTP & STDIO

Manual pip install

pip install marm-mcp-server
⁠Use this quick rule of thumb to choose your setup
  • Local HTTP/STDIO = fastest single-machine setup.
  • Docker HTTP = shared/always-on server (key required).
  • Docker STDIO = private containerized local use (no HTTP key).

Swarm / multi-agent note: The write queue is enabled by default to serialize memory writes through one worker. For shared HTTP deployments, use marm-memory start --profile swarm (200 RPM) or --profile swarm-max (600 RPM). --profile trusted disables rate limiting entirely for private deployments. STDIO is still best for private single-agent/local use. See Swarm & multi-agent presets⁠ for the full table.

Local pip HTTP

"agent" refers to claude, gemini, grok, qwen, or any MCP client. Codex uses --url instead of --transport to add MCP tools.

pip install marm-mcp-server
marm-memory start
# Stuck on client setup? Open a Q&A thread: https://github.com/Lyellr88/marm-memory/discussions
# Most agents use this --transport command
"agent" mcp add --transport http marm-memory http://localhost:8001/mcp
codex mcp add marm-memory --url http://localhost:8001/mcp

Default pip/local startup is zero-config: MARM binds to localhost and does not require a key unless you expose it with SERVER_HOST=0.0.0.0.

Local pip STDIO
pip install marm-mcp-server
python -m marm_mcp_server.server_stdio
# most agents use this --transport command
"agent" mcp add --transport stdio marm-memory-stdio marm-mcp-stdio
codex mcp add marm-memory-stdio -- marm-mcp-stdio

Replace marm-mcp-stdio with python -m marm_mcp_server.server_stdio if using a virtualenv or a path-based setup. Works with Claude Code, Cursor, VS Code, Qwen, and Gemini CLI. STDIO stays a single local process with no port and no API key, and exposes the same 14 tools as HTTP.

Local Python swarm modes (HTTP & STDIO)

Use HTTP when multiple agents need to share one live MARM server. STDIO is still best for private single-agent use because each client owns its own local process.

# HTTP shared server, normal multi-agent use
marm-memory start --profile swarm

# HTTP shared server, heavier private swarm
marm-memory start --profile swarm-max

# HTTP trusted private lab/server, rate limiting disabled
marm-memory start --profile trusted

# STDIO remains keyless/private and does not use swarm flags
marm-mcp-stdio

Docker HTTP (key required)

Docker HTTP requires an API key because it exposes MARM as a network server; STDIO stays local to the client process and does not need one.

If you installed MARM through pip, the product CLI can safely preview or run the same setup. It uses a loopback port by default, preserves ~/.marm, stores the generated key in ~/.marm/.env rather than shell history, and refuses to replace an existing container.

marm-memory docker command                 # preview the exact HTTP command
marm-memory docker run                     # create the managed HTTP container
marm-memory docker stdio-command           # print a Docker STDIO client command
marm-memory docker status
marm-memory docker logs --follow
marm-memory docker stop

# Optional: mount repositories read-only for code indexing.
marm-memory docker run --repo /absolute/path/to/repository

# Optional: preview or explicitly write a Compose configuration.
marm-memory docker compose
marm-memory docker compose --yes

The HTTP run, command, and compose commands accept the same operational flags:

For example:

# Shared local server with a custom data path and two repositories for indexing.
marm-memory docker command \
  --profile swarm \
  --data-dir /srv/marm-data \
  --repo /srv/projects/api \
  --repo /srv/projects/web

# Execute the reviewed command, pulling the image first.
marm-memory docker run --profile swarm --data-dir /srv/marm-data --pull
# Step 1: generate key (do not add < > around the key)
docker run --rm lyellr88/marm-mcp-server:latest --generate-key

# Step 2: run server
docker pull lyellr88/marm-mcp-server:latest
docker run -d --name marm-mcp-server \
  -p 127.0.0.1:8001:8001 \
  -e SERVER_HOST=0.0.0.0 \
  -e MARM_API_KEY=your-generated-key \
  -v ~/.marm:/home/marm/.marm \
  lyellr88/marm-mcp-server:latest

# Step 3: connect client
"agent" mcp add --transport http marm-memory http://localhost:8001/mcp --header "Authorization: Bearer your-generated-key"

# PowerShell: set this before starting/restarting Codex
$env:MARM_API_KEY="your-generated-key"
codex mcp add marm-memory --url http://localhost:8001/mcp --bearer-token-env-var MARM_API_KEY

# Quick auth smoke test
curl -i -H "Authorization: Bearer $env:MARM_API_KEY" http://127.0.0.1:8001/mcp
Docker HTTP swarm mode
# --swarm: write queue on, 200 RPM - recommended for multi-agent shared servers
docker run -d --name marm-mcp-server \
  -p 127.0.0.1:8001:8001 \
  -e SERVER_HOST=0.0.0.0 \
  -e MARM_API_KEY=your-generated-key \
  -v ~/.marm:/home/marm/.marm \
  lyellr88/marm-mcp-server:latest --swarm
Docker graph indexing: mount the repo

Docker graph tools run inside the container, so they cannot see host paths unless you mount them at docker run.

$env:MARM_API_KEY="test"

# The second -v line mounts your repo; adjust the host path to your project
docker run -d --name marm-mcp-server `
  -p 127.0.0.1:8001:8001 `
  -e SERVER_HOST=0.0.0.0 `
  -e MARM_API_KEY=$env:MARM_API_KEY `
  -v ~/.marm:/home/marm/.marm `
  -v C:\Users\lyell\Desktop\marm-memory:/workspace/marm-memory `
  lyellr88/marm-mcp-server:latest

Then index the container path, not the Windows host path:

marm_graph_index(repo_path="/workspace/marm-memory")

Graph tools must use the container path. Mounts cannot be added to an already-running container; stop and restart the container with the repo mount when you want Docker graph indexing.

Docker STDIO (no HTTP key)

Docker STDIO includes the same built-in marm-graph tools; no extra image or install step is required.

docker run --rm -i \
  -v ~/.marm:/home/marm/.marm \
  --entrypoint python \
  lyellr88/marm-mcp-server:latest \
  -m marm_mcp_server.server_stdio

⁠Complete MCP Tool Suite (14 Tools)

💡 Pro Tip: You don't need to manually call these tools! Just tell your AI agent what you want in natural language:

  • "Claude, log this session as 'Project Alpha' and add this conversation as 'database design discussion'"
  • "Remember this code snippet in your notebook for later"
  • "Search for what we discussed about authentication yesterday"

The AI agent will automatically use the appropriate tools. Manual tool access is available for power users who want direct control.

⁠🧠 Core Memory (7 tools)
ToolWhat it doesKey parameters
marm_smart_recallHybrid memory recall with an additive, bounded concept/code graph sidecar when a compatible graph existsquery, limit, session_name, search_all, detail=1/2/3, project, platform, exact_mode
marm_log_entryAdd structured session log entries; each entry is also embedded into semantic memory so marm_smart_recall can find itentry, session_name
marm_log_showDisplay all entries and sessions, with filteringsession_name
marm_deleteDelete a log session, log entry, or notebook entrytype, target, session_name, project, platform
marm_summaryCached, paste-ready session summaries with intelligent truncationsession_name
marm_notebookSession-scoped scratch pad plus promotion to a permanent, graph-linked docaction="add"|"use"|"show"|"status"|"clear"|"save", name, data, session_name, project, platform
marm_compactionAgent-assisted memory cleanup with a reviewable audit trailaction="status"|"candidates"|"review"|"stage"|"apply"|"discard"
⁠🕸️ Code Graph (5 tools)
ToolWhat it doesKey parameters
marm_graph_indexIndex a repo into the code-structure graph, check status, list projects, or turn automatic re-indexing on and offrepo_path, project, action
marm_code_lookupFind symbols, text patterns, or a symbol's source; use instead of grep/globkind="auto"|"symbol"|"text"|"snippet"
marm_graph_traceTrace call paths and data flow from a functiondirection, mode
marm_graph_architectureArchitecture overview: modules, node/edge breakdown, schemaproject
marm_graph_impactBlast radius of code changes: git diff → affected symbols + risksince, base_branch, depth
⁠🧩 Concept Graph (2 tools)
ToolWhat it doesKey parameters
marm_concept_buildRebuild the graph, or index memories stored before automatic indexing. New memories are indexed on their ownsession_name, project, or search_all=True (one required)
marm_concept_recallExplicitly query entities, relationships, and linked code symbolsquery, depth (1-5), direction, project, platform

Tag summary

Content type

Image

Digest

sha256:59679b865…

Size

293.3 MB

Last updated

about 22 hours ago

docker pull lyellr88/marm-mcp-server