Bridges multiple AI models and CLIs, enabling orchestrated workflows across Claude Code, Gemini CLI, Codex CLI, and other AI development tools.
10K+
18 Tools
Version 4.43 or later needs to be installed to add the server automatically
Use cases
About
Bridges multiple AI models and CLIs, enabling orchestrated workflows across Claude Code, Gemini CLI, Codex CLI, and other AI development tools.
| Attribute | Details |
|---|---|
| Docker Image | mcp/zen |
| Author | BeehiveInnovations |
| Repository | https://github.com/beehiveinnovations/zen-mcp-server |
| Attribute | Details |
|---|---|
| Dockerfile | https://github.com/beehiveinnovations/zen-mcp-server/blob/7afc7c1cc96e23992c8f105f960132c657883bb1/Dockerfile |
| Commit | 7afc7c1cc96e23992c8f105f960132c657883bb1 |
| Docker Image built by | Docker Inc. |
| Docker Scout Health Score | |
| Verify Signature | COSIGN_REPOSITORY=mcp/signatures cosign verify mcp/zen --key https://raw.githubusercontent.com/docker/keyring/refs/heads/main/public/mcp/latest.pub |
| Licence | Other |
| Tools provided by this Server | Short Description |
|---|---|
analyze | Performs comprehensive code analysis with systematic investigation and expert validation. |
apilookup | Use this tool automatically when you need current API/SDK documentation, latest version info, breaking changes, deprecations, migration guides, or official release notes. |
challenge | Prevents reflexive agreement by forcing critical thinking and reasoned analysis when a statement is challenged. |
chat | General chat and collaborative thinking partner for brainstorming, development discussion, getting second opinions, and exploring ideas. |
clink | Link a request to an external AI CLI (Gemini CLI, Qwen CLI, etc.) through PAL MCP to reuse their capabilities inside existing workflows. |
codereview | Performs systematic, step-by-step code review with expert validation. |
consensus | Builds multi-model consensus through systematic analysis and structured debate. |
debug | Performs systematic debugging and root cause analysis for any type of issue. |
docgen | Generates comprehensive code documentation with systematic analysis of functions, classes, and complexity. |
listmodels | Shows which AI model providers are configured, available model names, their aliases and capabilities. |
planner | Breaks down complex tasks through interactive, sequential planning with revision and branching capabilities. |
precommit | Validates git changes and repository state before committing with systematic analysis. |
refactor | Analyzes code for refactoring opportunities with systematic investigation. |
secaudit | Performs comprehensive security audit with systematic vulnerability assessment. |
testgen | Creates comprehensive test suites with edge case coverage for specific functions, classes, or modules. |
thinkdeep | Performs multi-stage investigation and reasoning for complex problem analysis. |
tracer | Performs systematic code tracing with modes for execution flow or dependency mapping. |
version | Get server version, configuration details, and list of available tools. |
analyzePerforms comprehensive code analysis with systematic investigation and expert validation. Use for architecture, performance, maintainability, and pattern analysis. Guides through structured code review and strategic planning.
| Parameters | Type | Description |
|---|---|---|
findings | string | Summary of discoveries from this step, including architectural patterns, tech stack assessment, scalability characteristics, performance implications, maintainability factors, and strategic improvement opportunities. IMPORTANT: Document both strengths (good patterns, solid architecture) and concerns (tech debt, overengineering, unnecessary complexity). In later steps, confirm or update past findings with additional evidence. |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | Set to true if you plan to continue the investigation with another step. False means you believe the analysis is complete and ready for expert validation. |
step | string | The analysis plan. Step 1: State your strategy, including how you will map the codebase structure, understand business logic, and assess code quality, performance implications, and architectural patterns. Later steps: Report findings and adapt the approach as new insights emerge. |
step_number | integer | The index of the current step in the analysis sequence, beginning at 1. Each step should build upon or revise the previous one. |
total_steps | integer | Your current estimate for how many steps will be needed to complete the analysis. Adjust as new findings emerge. |
analysis_type | stringoptional | Type of analysis to perform (architecture, performance, security, quality, general) |
confidence | stringoptional | Your confidence in the analysis: exploring, low, medium, high, very_high, almost_certain, or certain. 'certain' indicates the analysis is complete and ready for validation. |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | List all files examined (absolute paths). Include even ruled-out files to track exploration path. |
images | arrayoptional | Optional absolute paths to architecture diagrams or visual references that help with analysis context. |
issues_found | arrayoptional | Issues or concerns identified during analysis, each with severity level (critical, high, medium, low) |
output_format | stringoptional | How to format the output (summary, detailed, actionable) |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Subset of files_checked directly relevant to analysis findings (absolute paths). Include files with significant patterns, architectural decisions, or strategic improvement opportunities. |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
apilookupUse this tool automatically when you need current API/SDK documentation, latest version info, breaking changes, deprecations, migration guides, or official release notes. This tool searches authoritative sources (official docs, GitHub, package registries) to ensure up-to-date accuracy.
| Parameters | Type | Description |
|---|---|---|
prompt | string | The API, SDK, library, framework, or technology you need current documentation, version info, breaking changes, or migration guidance for. |
This tool is read-only. It does not modify its environment.
challengePrevents reflexive agreement by forcing critical thinking and reasoned analysis when a statement is challenged. Trigger automatically when a user critically questions, disagrees or appears to push back on earlier answers, and use it manually to sanity-check contentious claims.
| Parameters | Type | Description |
|---|---|---|
prompt | string | Statement to scrutinize. If you invoke challenge manually, strip the word 'challenge' and pass just the statement. Automatic invocations send the full user message as-is; do not modify it. |
This tool is read-only. It does not modify its environment.
chatGeneral chat and collaborative thinking partner for brainstorming, development discussion, getting second opinions, and exploring ideas. Use for ideas, validations, questions, and thoughtful explanations.
| Parameters | Type | Description |
|---|---|---|
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
prompt | string | Your question or idea for collaborative thinking to be sent to the external model. Provide detailed context, including your goal, what you've tried, and any specific challenges. WARNING: Large inline code must NOT be shared in prompt. Provide full-path to files on disk as separate parameter. |
working_directory_absolute_path | string | Absolute path to an existing directory where generated code artifacts can be saved. |
absolute_file_paths | arrayoptional | Full, absolute file paths to relevant code in order to share with external model |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
images | arrayoptional | Image paths (absolute) or base64 strings for optional visual context. |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
clinkLink a request to an external AI CLI (Gemini CLI, Qwen CLI, etc.) through PAL MCP to reuse their capabilities inside existing workflows.
| Parameters | Type | Description |
|---|---|---|
cli_name | string | Configured CLI client name (from conf/cli_clients). Available: claude, codex, gemini |
prompt | string | User request forwarded to the CLI (conversation context is pre-applied). |
absolute_file_paths | arrayoptional | Full paths to relevant code |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
images | arrayoptional | Optional absolute image paths or base64 blobs for visual context. |
role | stringoptional | Optional role preset defined for the selected CLI (defaults to 'default'). Roles per CLI: claude: codereviewer, default, planner; codex: codereviewer, default, planner; gemini: codereviewer, default, planner |
This tool is read-only. It does not modify its environment.
codereviewPerforms systematic, step-by-step code review with expert validation. Use for comprehensive analysis covering quality, security, performance, and architecture. Guides through structured investigation to ensure thoroughness.
| Parameters | Type | Description |
|---|---|---|
findings | string | Capture findings (positive and negative) across quality, security, performance, and architecture; update each step. |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | True when another review step follows. External validation: step 1 → True, step 2 → False. Internal validation: set False immediately. Apply the same rule on continuation flows. |
step | string | Review narrative. Step 1: outline the review strategy. Later steps: report findings. MUST cover quality, security, performance, and architecture. Reference code via relevant_files; avoid dumping large snippets. |
step_number | integer | Current review step (starts at 1) – each step should build on the last. |
total_steps | integer | Number of review steps planned. External validation: two steps (analysis + summary). Internal validation: one step. Use the same limits when continuing an existing review via continuation_id. |
confidence | stringoptional | Confidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed) |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | Absolute paths of every file reviewed, including those ruled out. |
focus_on | stringoptional | Optional note on areas to emphasise (e.g. 'threading', 'auth flow'). |
hypothesis | stringoptional | Current theory about issue/goal based on work |
images | arrayoptional | Optional diagram or screenshot paths that clarify review context. |
issues_found | arrayoptional | Issues with severity (critical/high/medium/low) and descriptions. |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Step 1: list all files/dirs under review. Must be absolute full non-abbreviated paths. Final step: narrow to files tied to key findings. |
review_type | stringoptional | Review focus: full, security, performance, or quick. |
review_validation_type | stringoptional | Set 'external' (default) for expert follow-up or 'internal' for local-only review. |
severity_filter | stringoptional | Lowest severity to include when reporting issues (critical/high/medium/low/all). |
standards | stringoptional | Coding standards or style guides to enforce. |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
consensusBuilds multi-model consensus through systematic analysis and structured debate. Use for complex decisions, architectural choices, feature proposals, and technology evaluations. Consults multiple models with different stances to synthesize comprehensive recommendations.
| Parameters | Type | Description |
|---|---|---|
findings | string | Step 1: your independent analysis for later synthesis (not shared with other models). Steps 2+: summarize the newest model response. |
next_step_required | boolean | True if more model consultations remain; set false when ready to synthesize. |
step | string | Consensus prompt. Step 1: write the exact proposal/question every model will see (use 'Evaluate…', not meta commentary). Steps 2+: capture internal notes about the latest model response—these notes are NOT sent to other models. |
step_number | integer | Current step index (starts at 1). Step 1 is your analysis; steps 2+ handle each model response. |
total_steps | integer | Total steps = number of models consulted plus the final synthesis step. |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
current_model_index | integeroptional | 0-based index of the next model to consult (managed internally). |
images | arrayoptional | Optional absolute image paths or base64 references that add helpful visual context. |
model_responses | arrayoptional | Internal log of responses gathered so far. |
models | arrayoptional | User-specified roster of models to consult (provide at least two entries). User-specified list of models to consult (provide at least two entries). Each entry may include model, stance (for/against/neutral), and stance_prompt. Each (model, stance) pair must be unique, e.g. [{'model':'gpt5','stance':'for'}, {'model':'pro','stance':'against'}]. When the user names a model, you MUST use that exact value or report the provider error—never swap in another option. Use the listmodels tool for the full roster. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
relevant_files | arrayoptional | Optional supporting files that help the consensus analysis. Must be absolute full, non-abbreviated paths. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
debugPerforms systematic debugging and root cause analysis for any type of issue. Use for complex bugs, mysterious errors, performance issues, race conditions, memory leaks, and integration problems. Guides through structured investigation with hypothesis testing and expert analysis.
| Parameters | Type | Description |
|---|---|---|
findings | string | Discoveries: clues, code/log evidence, disproven theories. Be specific. If no bug found, document clearly as valid. |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | True if you plan to continue the investigation with another step. False means root cause is known or investigation is complete. IMPORTANT: When continuation_id is provided (continuing a previous conversation), set this to False to immediately proceed with expert analysis. |
step | string | Investigation step. Step 1: State issue+direction. Symptoms misleading; 'no bug' valid. Trace dependencies, verify hypotheses. Use relevant_files for code; this for text only. |
step_number | integer | Current step index (starts at 1). Build upon previous steps. |
total_steps | integer | Estimated total steps needed to complete the investigation. Adjust as new findings emerge. IMPORTANT: When continuation_id is provided (continuing a previous conversation), set this to 1 as we're not starting a new multi-step investigation. |
confidence | stringoptional | Your confidence in the hypothesis: exploring (starting out), low (early idea), medium (some evidence), high (strong evidence), very_high (very strong evidence), almost_certain (nearly confirmed), certain (100% confidence - root cause and fix are both confirmed locally with no need for external validation). WARNING: Do NOT use 'certain' unless the issue can be fully resolved with a fix, use 'very_high' or 'almost_certain' instead when not 100% sure. Using 'certain' means you have ABSOLUTE confidence locally and PREVENTS external model validation. |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | All examined files (absolute paths), including ruled-out ones. |
hypothesis | stringoptional | Concrete root cause theory from evidence. Can revise. Valid: 'No bug found - user misunderstanding' or 'Symptoms unrelated to code' if supported. |
images | arrayoptional | Optional screenshots/visuals clarifying issue (absolute paths). |
issues_found | arrayoptional | Issues identified with severity levels during work |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Files directly relevant to issue (absolute paths). Cause, trigger, or manifestation locations. |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
docgenGenerates comprehensive code documentation with systematic analysis of functions, classes, and complexity. Use for documentation generation, code analysis, complexity assessment, and API documentation. Analyzes code structure and patterns to create thorough documentation.
| Parameters | Type | Description |
|---|---|---|
comments_on_complex_logic | boolean | True (default) to add inline comments around non-obvious logic. |
document_complexity | boolean | Include algorithmic complexity (Big O) analysis when True (default). |
document_flow | boolean | Include call flow/dependency notes when True (default). |
findings | string | Important findings, evidence and insights discovered in this step |
next_step_required | boolean | Whether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch. |
num_files_documented | integer | Count of files finished so far. Increment only when a file is fully documented. |
step | string | Current work step content and findings from your overall work |
step_number | integer | Current step number in work sequence (starts at 1) |
total_files_to_document | integer | Total files identified in discovery; completion requires matching this count. |
total_steps | integer | Estimated total steps needed to complete work |
update_existing | boolean | True (default) to polish inaccurate or outdated docs instead of leaving them untouched. |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
issues_found | arrayoptional | Issues identified with severity levels during work |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Files identified as relevant to issue/goal (FULL absolute paths to real files/folders - DO NOT SHORTEN) |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
listmodelsShows which AI model providers are configured, available model names, their aliases and capabilities.
plannerBreaks down complex tasks through interactive, sequential planning with revision and branching capabilities. Use for complex project planning, system design, migration strategies, and architectural decisions. Builds plans incrementally with deep reflection for complex scenarios.
| Parameters | Type | Description |
|---|---|---|
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | Whether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch. |
step | string | Planning content for this step. Step 1: describe the task, problem and scope. Later steps: capture updates, revisions, branches, or open questions that shape the plan. |
step_number | integer | Current step number in work sequence (starts at 1) |
total_steps | integer | Estimated total steps needed to complete work |
branch_from_step | integeroptional | If branching, the step number that this branch starts from. |
branch_id | stringoptional | Name for this branch (e.g. 'approach-A', 'migration-path'). |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
is_branch_point | booleanoptional | True when this step creates a new branch to explore an alternative path. |
is_step_revision | booleanoptional | Set true when you are replacing a previously recorded step. |
more_steps_needed | booleanoptional | True when you now expect to add additional steps beyond the prior estimate. |
revises_step_number | integeroptional | Step number being replaced when revising. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
precommitValidates git changes and repository state before committing with systematic analysis. Use for multi-repository validation, security review, change impact assessment, and completeness verification. Guides through structured investigation with expert analysis.
| Parameters | Type | Description |
|---|---|---|
findings | string | Record git diff insights, risks, missing tests, security concerns, and positives; update previous notes as you go. |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | True to continue with another step, False when validation is complete. CRITICAL: If total_steps>=3 or when precommit_type = external, set to True until the final step. When continuation_id is provided: Follow the same validation rules based on precommit_type. |
step | string | Step 1: outline how you'll validate the git changes. Later steps: report findings. Review diffs and impacts, use relevant_files, and avoid pasting large snippets. |
step_number | integer | Current pre-commit step number (starts at 1). |
total_steps | integer | Planned number of validation steps. External validation: use at most three (analysis → follow-ups → summary). Internal validation: a single step. Honour these limits when resuming via continuation_id. |
compare_to | stringoptional | Optional git ref (branch/tag/commit) to diff against; falls back to staged/unstaged changes. |
confidence | stringoptional | Confidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed) |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | Absolute paths for every file examined, including ruled-out candidates. |
focus_on | stringoptional | Optional emphasis areas such as security, performance, or test coverage. |
hypothesis | stringoptional | Current theory about issue/goal based on work |
images | arrayoptional | Optional absolute paths to screenshots or diagrams that aid validation. |
include_staged | booleanoptional | Whether to inspect staged changes (ignored when compare_to is set). |
include_unstaged | booleanoptional | Whether to inspect unstaged changes (ignored when compare_to is set). |
issues_found | arrayoptional | List issues with severity (critical/high/medium/low) plus descriptions (bugs, security, performance, coverage). |
path | stringoptional | Absolute path to the repository root. Required in step 1. |
precommit_type | stringoptional | 'external' (default, triggers expert model) or 'internal' (local-only validation). |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Absolute paths of files involved in the change or validation (code, configs, tests, docs). Must be absolute full non-abbreviated paths. |
severity_filter | stringoptional | Lowest severity to include when reporting issues. |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
refactorAnalyzes code for refactoring opportunities with systematic investigation. Use for code smell detection, decomposition planning, modernization, and maintainability improvements. Guides through structured analysis with expert validation.
| Parameters | Type | Description |
|---|---|---|
findings | string | Summary of discoveries from this step, including code smells and opportunities for decomposition, modernization, or organization. Document both strengths and weaknesses. In later steps, confirm or update past findings. |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | Set to true if you plan to continue the investigation with another step. False means you believe the refactoring analysis is complete and ready for expert validation. |
step | string | The refactoring plan. Step 1: State strategy. Later steps: Report findings. CRITICAL: Examine code for smells, and opportunities for decomposition, modernization, and organization. Use 'relevant_files' for code. FORBIDDEN: Large code snippets. |
step_number | integer | The index of the current step in the refactoring investigation sequence, beginning at 1. Each step should build upon or revise the previous one. |
total_steps | integer | Your current estimate for how many steps will be needed to complete the refactoring investigation. Adjust as new opportunities emerge. |
confidence | stringoptional | Your confidence in refactoring analysis: exploring (starting), incomplete (significant work remaining), partial (some opportunities found, more analysis needed), complete (comprehensive analysis finished, all major opportunities identified). WARNING: Use 'complete' ONLY when fully analyzed and can provide recommendations without expert help. 'complete' PREVENTS expert validation. Use 'partial' for large files or uncertain analysis. |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | List all files examined (absolute paths). Include even ruled-out files to track exploration path. |
focus_areas | arrayoptional | Specific areas to focus on (e.g., 'performance', 'readability', 'maintainability', 'security') |
hypothesis | stringoptional | Current theory about issue/goal based on work |
images | arrayoptional | Optional list of absolute paths to architecture diagrams, UI mockups, design documents, or visual references that help with refactoring context. Only include if they materially assist understanding or assessment. |
issues_found | arrayoptional | Refactoring opportunities as dictionaries with 'severity' (critical/high/medium/low), 'type' (codesmells/decompose/modernize/organization), and 'description'. Include all improvement opportunities found. |
refactor_type | stringoptional | Type of refactoring analysis to perform (codesmells, decompose, modernize, organization) |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Subset of files_checked with code requiring refactoring (absolute paths). Include files with code smells, decomposition needs, or improvement opportunities. |
style_guide_examples | arrayoptional | Optional existing code files to use as style/pattern reference (must be FULL absolute paths to real files / folders - DO NOT SHORTEN). These files represent the target coding style and patterns for the project. |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
secauditPerforms comprehensive security audit with systematic vulnerability assessment. Use for OWASP Top 10 analysis, compliance evaluation, threat modeling, and security architecture review. Guides through structured security investigation with expert validation.
| Parameters | Type | Description |
|---|---|---|
findings | string | Summarize vulnerabilities, auth issues, validation gaps, compliance notes, and positives; update prior findings as needed. |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | True while additional threat analysis remains; set False once you are ready to hand off for validation. |
step | string | Step 1: outline the audit strategy (OWASP Top 10, auth, validation, etc.). Later steps: report findings. MANDATORY: use relevant_files for code references and avoid large snippets. |
step_number | integer | Current security-audit step number (starts at 1). |
total_steps | integer | Expected number of audit steps; adjust as new risks surface. |
audit_focus | stringoptional | Primary focus area: owasp, compliance, infrastructure, dependencies, or comprehensive. |
compliance_requirements | arrayoptional | Applicable compliance frameworks or standards (SOC2, PCI DSS, HIPAA, GDPR, ISO 27001, NIST, etc.). |
confidence | stringoptional | exploring/low/medium/high/very_high/almost_certain/certain. 'certain' blocks external validation—use only when fully complete. |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | Absolute paths for every file inspected, including rejected candidates. |
hypothesis | stringoptional | Current theory about issue/goal based on work |
images | arrayoptional | Optional absolute paths to diagrams or threat models that inform the audit. |
issues_found | arrayoptional | Security issues with severity (critical/high/medium/low) and descriptions (vulns, auth flaws, injection, crypto, config). |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Absolute paths for security-relevant files (auth modules, configs, sensitive code). |
security_scope | stringoptional | Security context (web, mobile, API, cloud, etc.) including stack, user types, data sensitivity, and threat landscape. |
severity_filter | stringoptional | Minimum severity to include when reporting security issues. |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
threat_level | stringoptional | Assess the threat level: low (internal/low-risk), medium (customer-facing/business data), high (regulated or sensitive), critical (financial/healthcare/PII). |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
testgenCreates comprehensive test suites with edge case coverage for specific functions, classes, or modules. Analyzes code paths, identifies failure modes, and generates framework-specific tests. Be specific about scope - target particular components rather than testing everything.
| Parameters | Type | Description |
|---|---|---|
findings | string | Summarise functionality, critical paths, edge cases, boundary conditions, error handling, and existing test patterns. Cover both happy and failure paths. |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | True while more investigation or planning remains; set False when test planning is ready for expert validation. |
step | string | Test plan for this step. Step 1: outline how you'll analyse structure, business logic, critical paths, and edge cases. Later steps: record findings and new scenarios as they emerge. |
step_number | integer | Current test-generation step (starts at 1) — each step should build on prior work. |
total_steps | integer | Estimated number of steps needed for test planning; adjust as new scenarios appear. |
confidence | stringoptional | Indicate your current confidence in the test generation assessment. Use: 'exploring' (starting analysis), 'low' (early investigation), 'medium' (some patterns identified), 'high' (strong understanding), 'very_high' (very strong understanding), 'almost_certain' (nearly complete test plan), 'certain' (100% confidence - test plan is thoroughly complete and all test scenarios are identified with no need for external model validation). Do NOT use 'certain' unless the test generation analysis is comprehensively complete, use 'very_high' or 'almost_certain' instead if not 100% sure. Using 'certain' means you have complete confidence locally and prevents external model validation. |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | Absolute paths of every file examined, including those ruled out. |
hypothesis | stringoptional | Current theory about issue/goal based on work |
images | arrayoptional | Optional absolute paths to diagrams or visuals that clarify the system under test. |
issues_found | arrayoptional | Issues identified with severity levels during work |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Absolute paths of code that requires new or updated tests (implementation, dependencies, existing test fixtures). |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
thinkdeepPerforms multi-stage investigation and reasoning for complex problem analysis. Use for architecture decisions, complex bugs, performance challenges, and security analysis. Provides systematic hypothesis testing, evidence-based investigation, and expert validation.
| Parameters | Type | Description |
|---|---|---|
findings | string | Important findings, evidence and insights discovered in this step |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | Whether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch. |
step | string | Current work step content and findings from your overall work |
step_number | integer | Current step number in work sequence (starts at 1) |
total_steps | integer | Estimated total steps needed to complete work |
confidence | stringoptional | Confidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed) |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | List of files examined during this work step |
focus_areas | arrayoptional | Focus aspects (architecture, performance, security, etc.) |
hypothesis | stringoptional | Current theory about issue/goal based on work |
images | arrayoptional | Optional absolute image paths or base64 blobs for visual context. |
issues_found | arrayoptional | Issues identified with severity levels during work |
problem_context | stringoptional | Additional context about problem/goal. Be expressive. |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Files identified as relevant to issue/goal (FULL absolute paths to real files/folders - DO NOT SHORTEN) |
temperature | numberoptional | 0 = deterministic · 1 = creative. |
thinking_mode | stringoptional | Reasoning depth: minimal, low, medium, high, or max. |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
tracerPerforms systematic code tracing with modes for execution flow or dependency mapping. Use for method execution analysis, call chain tracing, dependency mapping, and architectural understanding. Supports precision mode (execution flow) and dependencies mode (structural relationships).
| Parameters | Type | Description |
|---|---|---|
findings | string | Important findings, evidence and insights discovered in this step |
model | string | Currently in auto model selection mode. CRITICAL: When the user names a model, you MUST use that exact name unless the server rejects it. If no model is provided, you may use the listmodels tool to review options and select an appropriate match. Top models: gpt-5.2 (score 100, 400K ctx, thinking, code-gen); gpt-5.1-codex (score 100, 400K ctx, thinking, code-gen); gemini-2.5-pro (score 100, 1.0M ctx, thinking, code-gen); gemini-3-pro-preview (score 100, 1.0M ctx, thinking, code-gen); gpt-5.2-pro (score 100, 400K ctx, thinking, code-gen); +26 more via listmodels. |
next_step_required | boolean | Whether another work step is needed. When false, aim to reduce total_steps to match step_number to avoid mismatch. |
step | string | Current work step content and findings from your overall work |
step_number | integer | Current step number in work sequence (starts at 1) |
target_description | string | Description of what to trace and WHY. Include context about what you're trying to understand or analyze. |
total_steps | integer | Estimated total steps needed to complete work |
trace_mode | string | Type of tracing: 'ask' (default - prompts user to choose mode), 'precision' (execution flow) or 'dependencies' (structural relationships) |
confidence | stringoptional | Confidence level: exploring (just starting), low (early investigation), medium (some evidence), high (strong evidence), very_high (comprehensive understanding), almost_certain (near complete confidence), certain (100% confidence locally - no external validation needed) |
continuation_id | stringoptional | Unique thread continuation ID for multi-turn conversations. Works across different tools. ALWAYS reuse the last continuation_id you were given—this preserves full conversation context, files, and findings so the agent can resume seamlessly. |
files_checked | arrayoptional | List of files examined during this work step |
images | arrayoptional | Optional paths to architecture diagrams or flow charts that help understand the tracing context. |
relevant_context | arrayoptional | Methods/functions identified as involved in the issue |
relevant_files | arrayoptional | Files identified as relevant to issue/goal (FULL absolute paths to real files/folders - DO NOT SHORTEN) |
use_assistant_model | booleanoptional | Use assistant model for expert analysis after workflow steps. False skips expert analysis, relies solely on your personal investigation. Defaults to True for comprehensive validation. |
This tool is read-only. It does not modify its environment.
versionGet server version, configuration details, and list of available tools.
{
"mcpServers": {
"zen": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"-e",
"OPENROUTER_API_KEY",
"-e",
"GEMINI_API_KEY",
"-e",
"OPENAI_API_KEY",
"-e",
"XAI_API_KEY",
"mcp/zen"
],
"env": {
"OPENROUTER_API_KEY": "sk-or-v1-****",
"GEMINI_API_KEY": "AIza****",
"OPENAI_API_KEY": "sk-proj-****",
"XAI_API_KEY": "xai-****"
}
}
}
}