Zen

Zen

Bridges multiple AI models and CLIs, enabling orchestrated workflows across Claude Code, Gemini CLI, Codex CLI, and other AI development tools.

10K+

18 Tools

Packaged by
Requires Secrets
Add to Docker Desktop

Version 4.43 or later needs to be installed to add the server automatically

Use cases

About

Zen MCP Server

Bridges multiple AI models and CLIs, enabling orchestrated workflows across Claude Code, Gemini CLI, Codex CLI, and other AI development tools.

What is an MCP Server?

MCP Info

Image Building Info

AttributeDetails
Dockerfilehttps://github.com/beehiveinnovations/zen-mcp-server/blob/7afc7c1cc96e23992c8f105f960132c657883bb1/Dockerfile
Commit7afc7c1cc96e23992c8f105f960132c657883bb1
Docker Image built byDocker Inc.
Docker Scout Health ScoreDocker Scout Health Score
Verify SignatureCOSIGN_REPOSITORY=mcp/signatures cosign verify mcp/zen --key https://raw.githubusercontent.com/docker/keyring/refs/heads/main/public/mcp/latest.pub
LicenceOther

Available Tools (18)

Tools provided by this ServerShort Description
analyzePerforms comprehensive code analysis with systematic investigation and expert validation.
apilookupUse this tool automatically when you need current API/SDK documentation, latest version info, breaking changes, deprecations, migration guides, or official release notes.
challengePrevents reflexive agreement by forcing critical thinking and reasoned analysis when a statement is challenged.
chatGeneral chat and collaborative thinking partner for brainstorming, development discussion, getting second opinions, and exploring ideas.
clinkLink a request to an external AI CLI (Gemini CLI, Qwen CLI, etc.) through PAL MCP to reuse their capabilities inside existing workflows.
codereviewPerforms systematic, step-by-step code review with expert validation.
consensusBuilds multi-model consensus through systematic analysis and structured debate.
debugPerforms systematic debugging and root cause analysis for any type of issue.
docgenGenerates comprehensive code documentation with systematic analysis of functions, classes, and complexity.
listmodelsShows which AI model providers are configured, available model names, their aliases and capabilities.
plannerBreaks down complex tasks through interactive, sequential planning with revision and branching capabilities.
precommitValidates git changes and repository state before committing with systematic analysis.
refactorAnalyzes code for refactoring opportunities with systematic investigation.
secauditPerforms comprehensive security audit with systematic vulnerability assessment.
testgenCreates comprehensive test suites with edge case coverage for specific functions, classes, or modules.
thinkdeepPerforms multi-stage investigation and reasoning for complex problem analysis.
tracerPerforms systematic code tracing with modes for execution flow or dependency mapping.
versionGet server version, configuration details, and list of available tools.

Tools Details

Tool: analyze

Performs comprehensive code analysis with systematic investigation and expert validation. Use for architecture, performance, maintainability, and pattern analysis. Guides through structured code review and strategic planning.

ParametersTypeDescription
findingsstringSummary of discoveries from this step, including architectural patterns, tech stack assessment, scalability characteristics, performance implications, maintainability factors, and strategic improvement opportunities. IMPORTANT: Document both strengths (good patterns, solid architecture) and concerns (tech debt, overengineering, unnecessary complexity). In later steps, confirm or update past findings with additional evidence.
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanSet to true if you plan to continue the investigation with another step. False means you believe the analysis is complete and ready for expert validation.
stepstringThe analysis plan. Step 1: State your strategy, including how you will map the codebase structure, understand business logic, and assess code quality, performance implications, and architectural patterns. Later steps: Report findings and adapt the approach as new insights emerge.
step_numberintegerThe index of the current step in the analysis sequence, beginning at 1. Each step should build upon or revise the previous one.
total_stepsintegerYour current estimate for how many steps will be needed to complete the analysis. Adjust as new findings emerge.
analysis_typestringoptionalType of analysis to perform (architecture, performance, security, quality, general)
confidencestringoptionalYour confidence in the analysis: exploring, low, medium, high, very_high, almost_certain, or certain. 'certain' indicates the analysis is complete and ready for validation.
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalList all files examined (absolute paths). Include even ruled-out files to track exploration path.
imagesarrayoptionalOptional absolute paths to architecture diagrams or visual references that help with analysis context.
issues_foundarrayoptionalIssues or concerns identified during analysis, each with severity level (critical, high, medium, low)
output_formatstringoptionalHow to format the output (summary, detailed, actionable)
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalSubset of files_checked directly relevant to analysis findings (absolute paths). Include files with significant patterns, architectural decisions, or strategic improvement opportunities.
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: apilookup

Use this tool automatically when you need current API/SDK documentation, latest version info, breaking changes, deprecations, migration guides, or official release notes. This tool searches authoritative sources (official docs, GitHub, package registries) to ensure up-to-date accuracy.

ParametersTypeDescription
promptstringThe API, SDK, library, framework, or technology you need current documentation, version info, breaking changes, or migration guidance for.

This tool is read-only. It does not modify its environment.


Tool: challenge

Prevents reflexive agreement by forcing critical thinking and reasoned analysis when a statement is challenged. Trigger automatically when a user critically questions, disagrees or appears to push back on earlier answers, and use it manually to sanity-check contentious claims.

ParametersTypeDescription
promptstringStatement to scrutinize. If you invoke challenge manually, strip the word 'challenge' and pass just the statement. Automatic invocations send the full user message as-is; do not modify it.

This tool is read-only. It does not modify its environment.


Tool: chat

General chat and collaborative thinking partner for brainstorming, development discussion, getting second opinions, and exploring ideas. Use for ideas, validations, questions, and thoughtful explanations.

ParametersTypeDescription
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
promptstringYour question or idea for collaborative thinking to be sent to the external model. Provide detailed context, including your goal, what you've tried, and any specific challenges. WARNING: Large inline code must NOT be shared in prompt. Provide full-path to files on disk as separate parameter.
working_directory_absolute_pathstringAbsolute path to an existing directory where generated code artifacts can be saved.
absolute_file_pathsarrayoptionalFull, absolute file paths to relevant code in order to share with external model
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
imagesarrayoptionalImage paths (absolute) or base64 strings for optional visual context.
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.

Tool: clink

Link a request to an external AI CLI (Gemini CLI, Qwen CLI, etc.) through PAL MCP to reuse their capabilities inside existing workflows.

ParametersTypeDescription
cli_namestringConfigured CLI client name (from conf/cli_clients). Available: claude, codex, gemini
promptstringUser request forwarded to the CLI (conversation context is pre-applied).
absolute_file_pathsarrayoptionalFull paths to relevant code
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
imagesarrayoptionalOptional absolute image paths or base64 blobs for visual context.
rolestringoptionalOptional role preset defined for the selected CLI (defaults to 'default'). Roles per CLI: claude: codereviewer, default, planner; codex: codereviewer, default, planner; gemini: codereviewer, default, planner

This tool is read-only. It does not modify its environment.


Tool: codereview

Performs systematic, step-by-step code review with expert validation. Use for comprehensive analysis covering quality, security, performance, and architecture. Guides through structured investigation to ensure thoroughness.

ParametersTypeDescription
findingsstringCapture findings (positive and negative) across quality, security, performance, and architecture; update each step.
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanTrue when another review step follows. External validation: step 1 → True, step 2 → False. Internal validation: set False immediately. Apply the same rule on continuation flows.
stepstringReview narrative. Step 1: outline the review strategy. Later steps: report findings. MUST cover quality, security, performance, and architecture. Reference code via relevant_files; avoid dumping large snippets.
step_numberintegerCurrent review step (starts at 1) – each step should build on the last.
total_stepsintegerNumber of review steps planned. External validation: two steps (analysis + summary). Internal validation: one step. Use the same limits when continuing an existing review via continuation_id.
confidencestringoptionalConfidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed)
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalAbsolute paths of every file reviewed, including those ruled out.
focus_onstringoptionalOptional note on areas to emphasise (e.g. 'threading', 'auth flow').
hypothesisstringoptionalCurrent theory about issue/goal based on work
imagesarrayoptionalOptional diagram or screenshot paths that clarify review context.
issues_foundarrayoptionalIssues with severity (critical/high/medium/low) and descriptions.
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalStep 1: list all files/dirs under review. Must be absolute full non-abbreviated paths. Final step: narrow to files tied to key findings.
review_typestringoptionalReview focus: full, security, performance, or quick.
review_validation_typestringoptionalSet 'external' (default) for expert follow-up or 'internal' for local-only review.
severity_filterstringoptionalLowest severity to include when reporting issues (critical/high/medium/low/all).
standardsstringoptionalCoding standards or style guides to enforce.
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: consensus

Builds multi-model consensus through systematic analysis and structured debate. Use for complex decisions, architectural choices, feature proposals, and technology evaluations. Consults multiple models with different stances to synthesize comprehensive recommendations.

ParametersTypeDescription
findingsstringStep 1: your independent analysis for later synthesis (not shared with other models). Steps 2+: summarize the newest model response.
next_step_requiredbooleanTrue if more model consultations remain; set false when ready to synthesize.
stepstringConsensus prompt. Step 1: write the exact proposal/question every model will see (use 'Evaluate…', not meta commentary). Steps 2+: capture internal notes about the latest model response—these notes are NOT sent to other models.
step_numberintegerCurrent step index (starts at 1). Step 1 is your analysis; steps 2+ handle each model response.
total_stepsintegerTotal steps = number of models consulted plus the final synthesis step.
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
current_model_indexintegeroptional0-based index of the next model to consult (managed internally).
imagesarrayoptionalOptional absolute image paths or base64 references that add helpful visual context.
model_responsesarrayoptionalInternal log of responses gathered so far.
modelsarrayoptionalUser-specified roster of models to consult (provide at least two entries). User-specified list of models to consult (provide at least two entries). Each entry may include model, stance (for/against/neutral), and stance_prompt. Each (model, stance) pair must be unique, e.g. [{'model':'gpt5','stance':'for'}, {'model':'pro','stance':'against'}]. When the user names a model, you MUST use that exact value or report the provider error—never swap in another option. Use the listmodels tool for the full roster. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
relevant_filesarrayoptionalOptional supporting files that help the consensus analysis. Must be absolute full, non-abbreviated paths.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: debug

Performs systematic debugging and root cause analysis for any type of issue. Use for complex bugs, mysterious errors, performance issues, race conditions, memory leaks, and integration problems. Guides through structured investigation with hypothesis testing and expert analysis.

ParametersTypeDescription
findingsstringDiscoveries: clues, code/log evidence, disproven theories. Be specific. If no bug found, document clearly as valid.
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanTrue if you plan to continue the investigation with another step. False means root cause is known or investigation is complete. IMPORTANT: When continuation_id is provided (continuing a previous conversation), set this to False to immediately proceed with expert analysis.
stepstringInvestigation step. Step 1: State issue+direction. Symptoms misleading; 'no bug' valid. Trace dependencies, verify hypotheses. Use relevant_files for code; this for text only.
step_numberintegerCurrent step index (starts at 1). Build upon previous steps.
total_stepsintegerEstimated total steps needed to complete the investigation. Adjust as new findings emerge. IMPORTANT: When continuation_id is provided (continuing a previous conversation), set this to 1 as we're not starting a new multi-step investigation.
confidencestringoptionalYour confidence in the hypothesis: exploring (starting out), low (early idea), medium (some evidence), high (strong evidence), very_high (very strong evidence), almost_certain (nearly confirmed), certain (100% confidence - root cause and fix are both confirmed locally with no need for external validation). WARNING: Do NOT use 'certain' unless the issue can be fully resolved with a fix, use 'very_high' or 'almost_certain' instead when not 100% sure. Using 'certain' means you have ABSOLUTE confidence locally and PREVENTS external model validation.
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalAll examined files (absolute paths), including ruled-out ones.
hypothesisstringoptionalConcrete root cause theory from evidence. Can revise. Valid: 'No bug found - user misunderstanding' or 'Symptoms unrelated to code' if supported.
imagesarrayoptionalOptional screenshots/visuals clarifying issue (absolute paths).
issues_foundarrayoptionalIssues identified with severity levels during work
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalFiles directly relevant to issue (absolute paths). Cause, trigger, or manifestation locations.
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: docgen

Generates comprehensive code documentation with systematic analysis of functions, classes, and complexity. Use for documentation generation, code analysis, complexity assessment, and API documentation. Analyzes code structure and patterns to create thorough documentation.

ParametersTypeDescription
comments_on_complex_logicbooleanTrue (default) to add inline comments around non-obvious logic.
document_complexitybooleanInclude algorithmic complexity (Big O) analysis when True (default).
document_flowbooleanInclude call flow/dependency notes when True (default).
findingsstringImportant findings, evidence and insights discovered in this step
next_step_requiredbooleanWhether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch.
num_files_documentedintegerCount of files finished so far. Increment only when a file is fully documented.
stepstringCurrent work step content and findings from your overall work
step_numberintegerCurrent step number in work sequence (starts at 1)
total_files_to_documentintegerTotal files identified in discovery; completion requires matching this count.
total_stepsintegerEstimated total steps needed to complete work
update_existingbooleanTrue (default) to polish inaccurate or outdated docs instead of leaving them untouched.
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
issues_foundarrayoptionalIssues identified with severity levels during work
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalFiles identified as relevant to issue/goal (FULL absolute paths to real files/folders - DO NOT SHORTEN)
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: listmodels

Shows which AI model providers are configured, available model names, their aliases and capabilities.

Tool: planner

Breaks down complex tasks through interactive, sequential planning with revision and branching capabilities. Use for complex project planning, system design, migration strategies, and architectural decisions. Builds plans incrementally with deep reflection for complex scenarios.

ParametersTypeDescription
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanWhether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch.
stepstringPlanning content for this step. Step 1: describe the task, problem and scope. Later steps: capture updates, revisions, branches, or open questions that shape the plan.
step_numberintegerCurrent step number in work sequence (starts at 1)
total_stepsintegerEstimated total steps needed to complete work
branch_from_stepintegeroptionalIf branching, the step number that this branch starts from.
branch_idstringoptionalName for this branch (e.g. 'approach-A', 'migration-path').
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
is_branch_pointbooleanoptionalTrue when this step creates a new branch to explore an alternative path.
is_step_revisionbooleanoptionalSet true when you are replacing a previously recorded step.
more_steps_neededbooleanoptionalTrue when you now expect to add additional steps beyond the prior estimate.
revises_step_numberintegeroptionalStep number being replaced when revising.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: precommit

Validates git changes and repository state before committing with systematic analysis. Use for multi-repository validation, security review, change impact assessment, and completeness verification. Guides through structured investigation with expert analysis.

ParametersTypeDescription
findingsstringRecord git diff insights, risks, missing tests, security concerns, and positives; update previous notes as you go.
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanTrue to continue with another step, False when validation is complete. CRITICAL: If total_steps>=3 or when precommit_type = external, set to True until the final step. When continuation_id is provided: Follow the same validation rules based on precommit_type.
stepstringStep 1: outline how you'll validate the git changes. Later steps: report findings. Review diffs and impacts, use relevant_files, and avoid pasting large snippets.
step_numberintegerCurrent pre-commit step number (starts at 1).
total_stepsintegerPlanned number of validation steps. External validation: use at most three (analysis → follow-ups → summary). Internal validation: a single step. Honour these limits when resuming via continuation_id.
compare_tostringoptionalOptional git ref (branch/tag/commit) to diff against; falls back to staged/unstaged changes.
confidencestringoptionalConfidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed)
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalAbsolute paths for every file examined, including ruled-out candidates.
focus_onstringoptionalOptional emphasis areas such as security, performance, or test coverage.
hypothesisstringoptionalCurrent theory about issue/goal based on work
imagesarrayoptionalOptional absolute paths to screenshots or diagrams that aid validation.
include_stagedbooleanoptionalWhether to inspect staged changes (ignored when compare_to is set).
include_unstagedbooleanoptionalWhether to inspect unstaged changes (ignored when compare_to is set).
issues_foundarrayoptionalList issues with severity (critical/high/medium/low) plus descriptions (bugs, security, performance, coverage).
pathstringoptionalAbsolute path to the repository root. Required in step 1.
precommit_typestringoptional'external' (default, triggers expert model) or 'internal' (local-only validation).
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalAbsolute paths of files involved in the change or validation (code, configs, tests, docs). Must be absolute full non-abbreviated paths.
severity_filterstringoptionalLowest severity to include when reporting issues.
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: refactor

Analyzes code for refactoring opportunities with systematic investigation. Use for code smell detection, decomposition planning, modernization, and maintainability improvements. Guides through structured analysis with expert validation.

ParametersTypeDescription
findingsstringSummary of discoveries from this step, including code smells and opportunities for decomposition, modernization, or organization. Document both strengths and weaknesses. In later steps, confirm or update past findings.
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanSet to true if you plan to continue the investigation with another step. False means you believe the refactoring analysis is complete and ready for expert validation.
stepstringThe refactoring plan. Step 1: State strategy. Later steps: Report findings. CRITICAL: Examine code for smells, and opportunities for decomposition, modernization, and organization. Use 'relevant_files' for code. FORBIDDEN: Large code snippets.
step_numberintegerThe index of the current step in the refactoring investigation sequence, beginning at 1. Each step should build upon or revise the previous one.
total_stepsintegerYour current estimate for how many steps will be needed to complete the refactoring investigation. Adjust as new opportunities emerge.
confidencestringoptionalYour confidence in refactoring analysis: exploring (starting), incomplete (significant work remaining), partial (some opportunities found, more analysis needed), complete (comprehensive analysis finished, all major opportunities identified). WARNING: Use 'complete' ONLY when fully analyzed and can provide recommendations without expert help. 'complete' PREVENTS expert validation. Use 'partial' for large files or uncertain analysis.
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalList all files examined (absolute paths). Include even ruled-out files to track exploration path.
focus_areasarrayoptionalSpecific areas to focus on (e.g., 'performance', 'readability', 'maintainability', 'security')
hypothesisstringoptionalCurrent theory about issue/goal based on work
imagesarrayoptionalOptional list of absolute paths to architecture diagrams, UI mockups, design documents, or visual references that help with refactoring context. Only include if they materially assist understanding or assessment.
issues_foundarrayoptionalRefactoring opportunities as dictionaries with 'severity' (critical/high/medium/low), 'type' (codesmells/decompose/modernize/organization), and 'description'. Include all improvement opportunities found.
refactor_typestringoptionalType of refactoring analysis to perform (codesmells, decompose, modernize, organization)
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalSubset of files_checked with code requiring refactoring (absolute paths). Include files with code smells, decomposition needs, or improvement opportunities.
style_guide_examplesarrayoptionalOptional existing code files to use as style/pattern reference (must be FULL absolute paths to real files / folders - DO NOT SHORTEN). These files represent the target coding style and patterns for the project.
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: secaudit

Performs comprehensive security audit with systematic vulnerability assessment. Use for OWASP Top 10 analysis, compliance evaluation, threat modeling, and security architecture review. Guides through structured security investigation with expert validation.

ParametersTypeDescription
findingsstringSummarize vulnerabilities, auth issues, validation gaps, compliance notes, and positives; update prior findings as needed.
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanTrue while additional threat analysis remains; set False once you are ready to hand off for validation.
stepstringStep 1: outline the audit strategy (OWASP Top 10, auth, validation, etc.). Later steps: report findings. MANDATORY: use relevant_files for code references and avoid large snippets.
step_numberintegerCurrent security-audit step number (starts at 1).
total_stepsintegerExpected number of audit steps; adjust as new risks surface.
audit_focusstringoptionalPrimary focus area: owasp, compliance, infrastructure, dependencies, or comprehensive.
compliance_requirementsarrayoptionalApplicable compliance frameworks or standards (SOC2, PCI DSS, HIPAA, GDPR, ISO 27001, NIST, etc.).
confidencestringoptionalexploring/low/medium/high/very_high/almost_certain/certain. 'certain' blocks external validation—use only when fully complete.
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalAbsolute paths for every file inspected, including rejected candidates.
hypothesisstringoptionalCurrent theory about issue/goal based on work
imagesarrayoptionalOptional absolute paths to diagrams or threat models that inform the audit.
issues_foundarrayoptionalSecurity issues with severity (critical/high/medium/low) and descriptions (vulns, auth flaws, injection, crypto, config).
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalAbsolute paths for security-relevant files (auth modules, configs, sensitive code).
security_scopestringoptionalSecurity context (web, mobile, API, cloud, etc.) including stack, user types, data sensitivity, and threat landscape.
severity_filterstringoptionalMinimum severity to include when reporting security issues.
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
threat_levelstringoptionalAssess the threat level: low (internal/low-risk), medium (customer-facing/business data), high (regulated or sensitive), critical (financial/healthcare/PII).
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: testgen

Creates comprehensive test suites with edge case coverage for specific functions, classes, or modules. Analyzes code paths, identifies failure modes, and generates framework-specific tests. Be specific about scope - target particular components rather than testing everything.

ParametersTypeDescription
findingsstringSummarise functionality, critical paths, edge cases, boundary conditions, error handling, and existing test patterns. Cover both happy and failure paths.
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanTrue while more investigation or planning remains; set False when test planning is ready for expert validation.
stepstringTest plan for this step. Step 1: outline how you'll analyse structure, business logic, critical paths, and edge cases. Later steps: record findings and new scenarios as they emerge.
step_numberintegerCurrent test-generation step (starts at 1) — each step should build on prior work.
total_stepsintegerEstimated number of steps needed for test planning; adjust as new scenarios appear.
confidencestringoptionalIndicate your current confidence in the test generation assessment. Use: 'exploring' (starting analysis), 'low' (early investigation), 'medium' (some patterns identified), 'high' (strong understanding), 'very_high' (very strong understanding), 'almost_certain' (nearly complete test plan), 'certain' (100% confidence - test plan is thoroughly complete and all test scenarios are identified with no need for external model validation). Do NOT use 'certain' unless the test generation analysis is comprehensively complete, use 'very_high' or 'almost_certain' instead if not 100% sure. Using 'certain' means you have complete confidence locally and prevents external model validation.
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalAbsolute paths of every file examined, including those ruled out.
hypothesisstringoptionalCurrent theory about issue/goal based on work
imagesarrayoptionalOptional absolute paths to diagrams or visuals that clarify the system under test.
issues_foundarrayoptionalIssues identified with severity levels during work
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalAbsolute paths of code that requires new or updated tests (implementation, dependencies, existing test fixtures).
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: thinkdeep

Performs multi-stage investigation and reasoning for complex problem analysis. Use for architecture decisions, complex bugs, performance challenges, and security analysis. Provides systematic hypothesis testing, evidence-based investigation, and expert validation.

ParametersTypeDescription
findingsstringImportant findings, evidence and insights discovered in this step
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanWhether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch.
stepstringCurrent work step content and findings from your overall work
step_numberintegerCurrent step number in work sequence (starts at 1)
total_stepsintegerEstimated total steps needed to complete work
confidencestringoptionalConfidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed)
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalList of files examined during this work step
focus_areasarrayoptionalFocus aspects (architecture, performance, security, etc.)
hypothesisstringoptionalCurrent theory about issue/goal based on work
imagesarrayoptionalOptional absolute image paths or base64 blobs for visual context.
issues_foundarrayoptionalIssues identified with severity levels during work
problem_contextstringoptionalAdditional context about problem/goal. Be expressive.
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalFiles identified as relevant to issue/goal (FULL absolute paths to real files/folders - DO NOT SHORTEN)
temperaturenumberoptional0 = deterministic · 1 = creative.
thinking_modestringoptionalReasoning depth: minimal, low, medium, high, or max.
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: tracer

Performs systematic code tracing with modes for execution flow or dependency mapping. Use for method execution analysis, call chain tracing, dependency mapping, and architectural understanding. Supports precision mode (execution flow) and dependencies mode (structural relationships).

ParametersTypeDescription
findingsstringImportant findings, evidence and insights discovered in this step
modelstringCurrently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels.
next_step_requiredbooleanWhether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch.
stepstringCurrent work step content and findings from your overall work
step_numberintegerCurrent step number in work sequence (starts at 1)
target_descriptionstringDescription of what to trace and WHY. Include context about what you're trying to understand or analyze.
total_stepsintegerEstimated total steps needed to complete work
trace_modestringType of tracing: 'ask' (default - prompts user to choose mode), 'precision' (execution flow) or 'dependencies' (structural relationships)
confidencestringoptionalConfidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed)
continuation_idstringoptionalUnique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly.
files_checkedarrayoptionalList of files examined during this work step
imagesarrayoptionalOptional paths to architecture diagrams or flow charts that help understand the tracing context.
relevant_contextarrayoptionalMethods/functions identified as involved in the issue
relevant_filesarrayoptionalFiles identified as relevant to issue/goal (FULL absolute paths to real files/folders - DO NOT SHORTEN)
use_assistant_modelbooleanoptionalUse assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation.

This tool is read-only. It does not modify its environment.


Tool: version

Get server version, configuration details, and list of available tools.

Use this MCP Server

{
  "mcpServers": {
    "zen": {
      "command": "docker",
      "args": [
        "run",
        "-i",
        "--rm",
        "-e",
        "OPENROUTER_API_KEY",
        "-e",
        "GEMINI_API_KEY",
        "-e",
        "OPENAI_API_KEY",
        "-e",
        "XAI_API_KEY",
        "mcp/zen"
      ],
      "env": {
        "OPENROUTER_API_KEY": "sk-or-v1-****",
        "GEMINI_API_KEY": "AIza****",
        "OPENAI_API_KEY": "sk-proj-****",
        "XAI_API_KEY": "xai-****"
      }
    }
  }
}

Why is it safer to run MCP Servers with Docker?

Related servers