Sign inSign up

superbizon007/plan-judge-agent

By superbizon007

•Updated 5 months ago

A planning agent that uses a **dual-planner + judge loop**

Image
0

495

superbizon007/plan-judge-agent repository overview

⁠plan-judge-agent

A planning agent that uses a dual-planner + judge loop to produce a validated JSON execution plan from a user request and a list of available A2A agent cards.

Two LLMs (Planner A, Planner B) generate competing plan candidates each turn. A Judge LLM scores them, picks the winner, and both planners refine it. The loop exits when the score reaches the acceptance threshold or the maximum number of turns is reached.

⁠Architecture

validate_input → index_agents → generate_candidates (A + B, parallel)
  → judge_score → check_updates
    ├─ score ≥ threshold OR turn ≥ max_turns → validate_graph
    │                                            → collect_required_agents
    │                                              → write_report → validate_output → END
    └─ else → generate_candidates (loop)

⁠Quick Start

# Build
docker build --build-context a2a_agent=../../packages/a2a_agent -t plan-judge-agent:latest .

# Run
docker run -p 443:443 \
  -e PLANNER_A_URL=https://api.openai.com/v1 \
  -e PLANNER_A_API_KEY=sk-... \
  -e PLANNER_A_MODEL=gpt-4o \
  -e PLANNER_B_URL=https://api.openai.com/v1 \
  -e PLANNER_B_API_KEY=sk-... \
  -e PLANNER_B_MODEL=gpt-4o \
  -e JUDGE_URL=https://api.openai.com/v1 \
  -e JUDGE_API_KEY=sk-... \
  -e JUDGE_MODEL=gpt-4o \
  plan-judge-agent:latest

To use the same endpoint for all three roles, set all *_URL, *_API_KEY, and *_MODEL variables to the same values.

⁠Environment Variables

VariableDefaultDescription
PLANNER_A_URL—Base URL of the LLM API for Planner A
PLANNER_A_API_KEY—API key for Planner A
PLANNER_A_MODEL—Model name for Planner A
PLANNER_B_URL—Base URL of the LLM API for Planner B
PLANNER_B_API_KEY—API key for Planner B
PLANNER_B_MODEL—Model name for Planner B
JUDGE_URL—Base URL of the LLM API for the Judge
JUDGE_API_KEY—API key for the Judge
JUDGE_MODEL—Model name for the Judge
JUDGE_ACCEPTANCE_THRESHOLD80Judge score (0–100) at which the loop exits early
MAX_JUDGE_TURNS3Maximum judge loop iterations
API_KEY—Bearer token required on all incoming requests (optional)
PORT443HTTPS server port
TRAJECTORY_DIR—Directory for trajectory output; disabled if unset

⁠API

⁠Health check
GET /health
⁠Agent card (A2A)
GET /.well-known/agent.json
⁠A2A endpoint

Send a JSON task payload to the A2A endpoint. The input message content must be a JSON string:

{
  "request": "Build a competitive analysis report for product X",
  "context": {
    "project": "Product X launch",
    "constraints": ["must complete within 48h"],
    "deadline": "2026-03-21",
    "budget": "500 USD"
  },
  "available_agents": [
    {
      "name": "ResearchAgent",
      "description": "Searches the web and summarizes findings",
      "url": "https://research-agent.example.com/",
      "skills": [{"id": "search", "name": "Web Search", "description": "...", "tags": ["search"]}]
    }
  ]
}
⁠OpenAI-compatible endpoint
POST /v1/chat/completions

Send the JSON payload as the content of the last user message. Supports both streaming ("stream": true) and non-streaming responses.

⁠Output

{
  "plan_id": "550e8400-e29b-41d4-a716-446655440000",
  "summary": "...",
  "assumptions": ["..."],
  "required_agents": [
    {"name": "ResearchAgent", "url": "https://...", "role": "research", "price_per_turn": 0.0}
  ],
  "step_graph": {
    "is_dag": true,
    "entry_node_ids": ["S1"],
    "exit_node_ids": ["S3"],
    "nodes": [...],
    "edges": [...]
  },
  "execution_handoff": {"mode": "plan_only", "ready_for_orchestrator": true}
}

On failure:

{
  "error": {
    "code": "PLANNING_FAILED",
    "message": "...",
    "details": {}
  }
}

Error codes: VALIDATION_ERROR, PLANNING_FAILED, OUTPUT_SCHEMA_VIOLATION.

⁠Trajectory

When TRAJECTORY_DIR is set, each run writes:

$TRAJECTORY_DIR/
  20260313_181523/
    plan_final.json
    turn_0000/
      planner_a_request.json
      planner_a_response.json
      planner_b_request.json
      planner_b_response.json
      judge_request.json
      judge_response.json
    turn_0001/
      ...

⁠Development

# Install dependencies
uv sync --extra server

# Run tests (no env vars required)
uv run --extra server pytest tests/test_input_contract.py tests/test_graph_validators.py tests/test_judge_loop.py -v

# Run server locally
uv run --extra server python3 -m app.server

⁠MCP tools

The agent image exposes an MCP (Model Context Protocol) server on the same port as A2A at /mcp. Any MCP client — Claude Desktop, Cursor, Windsurf, ChatGPT Connectors, OpenAI Agents SDK — can list and invoke these tools using the pod's API_KEY as Bearer.

In production, Core proxies https://plan-judge-agent.agents.forfetch.ai/mcp → the worker pod's /mcp.

Tool (skill_id)Description
planReview and score a plan for correctness and feasibility

Input schemas are authored in apps/agents_mcp/app/seed.py. Skills without an explicit schema advertise a single {task: str} freeform parameter.

Tag summary

Content type

Image

Digest

sha256:6d7091f20…

Size

139 MB

Last updated

5 months ago

docker pull superbizon007/plan-judge-agent