Multi-agent system for Kubernetes orchestration, progressive delivery, and live debugging.
478
An open-source, extensible AI operations framework for Kubernetes — with built-in operators for Helm, GitOps, progressive delivery, and full-stack observability.
Managing Kubernetes in production usually means jumping between a half-dozen browser tabs and terminal windows: pulling container logs, reading events, inspecting Prometheus dashboards, checking ArgoCD sync status, verifying Traefik routes, and looking up Helm revision history.
Most AI tools for DevOps try to solve this by acting as generic chatbot wrappers around kubectl. They lack real context on your cluster, cannot verify whether their proposed commands actually fixed anything, and running them with write permissions feels like an unnecessary risk to production stability.
k8s-autopilot takes a fundamentally different approach. It is an open-source, extensible operations framework that connects an autonomous multi-agent runtime directly to your live infrastructure through the Model Context Protocol (MCP). Instead of generating unverified scripts, it investigates your cluster, drafts a clear step-by-step plan, asks for your confirmation before changing anything, executes the work, and independently verifies the result against concrete acceptance criteria.
k8s-autopilot is not a single hardcoded assistant. It is a modular framework built on LangGraph and the Deep Agents SDK. The core platform provides the orchestration layer, safety engine, verification loops, dynamic tool routing, and persistent memory.
User (Web UI / Chat / API)
│
▼
Central Supervisor Agent
│
┌────────────────────────┼────────────────────────┐
▼ ▼ ▼
Built-in Operators Marketplace Plugins Custom Sub-agents
(Helm, App, K8s, Obs) (Claude, Codex, OpsCode) (Dynamic Specialist Agents)
│ │ │
└────────────────────────┼────────────────────────┘
│
▼
Model Context Protocol (MCP)
(11 built-in stdio servers + external HTTP/SSE)
│
▼
Live Kubernetes Clusters & Infrastructure
The framework is designed so you decide what the agent can do. You can connect public or internal Git marketplaces and install plugins with a single click from the UI:
Every integration in k8s-autopilot is built on the open Model Context Protocol (MCP).
stdio, HTTP, or SSE directly from the Settings UI.Most agents consider a task "done" as soon as an API call returns a 200 OK. k8s-autopilot treats operational tasks as verifiable goals:
Safety is the core design priority of the framework. You control how much autonomy the agent has:
169.254.169.254).The agent remembers what it learns about your environments across conversations using AGENTS.md:
The web interface pairs the open Agent-to-Agent (A2A) communication protocol with Agent-to-User Interface (A2UI) visual blueprints:
Out of the box, k8s-autopilot ships with four domain operators that handle the most common Kubernetes workflows:
| Operator | Focus Area | Primary Capabilities |
|---|---|---|
| Kubernetes Operator | Core Cluster Operations | Diagnoses CrashLoopBackOff and OOMKilled pods; inspects container logs; safe pod exec and ephemeral debugging; scales deployments; audits RBAC roles and service accounts; switches multi-cluster contexts. |
| Helm Operator | Package Lifecycle | Discovers, installs, upgrades, and uninstalls releases; validates values against JSON schemas; dry-run preview; preserves values on upgrade; uses verified revision history for rollbacks; generates Helm charts and commits them to Git. |
| App Operator | GitOps & Progressive Delivery | Manages ArgoCD application lifecycle and sync debugging; converts Deployments to Argo Rollouts with zero downtime; executes canary and blue-green deployments with automated Prometheus metric analysis; configures Traefik edge routing, weighted splitting, and middlewares. |
| Observability Operator | Full-Stack Monitoring | Translates natural language into PromQL, LogQL, and TraceQL queries; cross-pillar incident correlation across Prometheus, Alertmanager, Loki, and Tempo; manages exporter lifecycles and OpenTelemetry auto-instrumentation pipelines; triage alerts and silences with blast-radius preview. |
The automated script checks prerequisites (Docker, Docker Compose, kubeconfig), sets up ~/.k8s-autopilot/, pulls the images, and starts the stack:
curl -LsSf https://raw.githubusercontent.com/talkops-ai/k8s-autopilot/main/scripts/install.sh | bash
Once started, open http://localhost:8888 in your browser.
You can launch the complete setup — including the PostgreSQL persistence backend and the TalkOps Web UI — using Docker Compose:
1. Create a docker-compose.yml file:
services:
postgres:
image: postgres:16-alpine
container_name: k8s-autopilot-postgres
restart: unless-stopped
environment:
- POSTGRES_USER=k8s_autopilot
- POSTGRES_PASSWORD=${POSTGRES_PASSWORD:-TalkOpsAutopilot2026SecureDB}
- POSTGRES_DB=k8s_autopilot
ports:
- "5432:5432"
volumes:
- postgres_data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U k8s_autopilot -d k8s_autopilot"]
interval: 5s
timeout: 5s
retries: 5
networks:
- k8s-autopilot-net
k8s-autopilot:
image: talkopsai/k8s-autopilot:latest
container_name: k8s-autopilot
ports:
- "10102:10102"
depends_on:
postgres:
condition: service_healthy
environment:
- POSTGRES_URI=postgresql://k8s_autopilot:${POSTGRES_PASSWORD:-TalkOpsAutopilot2026SecureDB}@postgres:5432/k8s_autopilot?sslmode=disable
- KUBECONFIG=/app/.kube/config
# Optional: set an API key here, or configure it via the Settings UI later
- GOOGLE_API_KEY=${GOOGLE_API_KEY}
- MODEL=${MODEL:-gemini-3.7-flash}
- REASONING_EFFORT=${REASONING_EFFORT:-medium}
volumes:
# Mount your kubeconfig so the agent can interact with your cluster
- ${HOME}/.kube/config:/app/.kube/config:ro
# Optional: mount local workspace for generated charts
- ./workspace/helm-charts:/app/workspace/helm-charts
restart: unless-stopped
networks:
- k8s-autopilot-net
talkops-ui:
image: talkopsai/talkops:latest
container_name: talkops-ui
environment:
- K8S_AGENT_URL=http://localhost:10102
- TALKOPS_ENABLE_LOGGING=false
- VITE_ENVIRONMENT=production
ports:
- "8888:80"
depends_on:
- k8s-autopilot
restart: unless-stopped
networks:
- k8s-autopilot-net
networks:
k8s-autopilot-net:
driver: bridge
volumes:
postgres_data:
2. Start the stack:
docker compose up -d
3. Connect and configure:
http://localhost:10102).docker run)To run the agent API directly with your local kubeconfig and an API key:
docker run -d \
--name k8s-autopilot \
-p 10102:10102 \
-e GOOGLE_API_KEY="your_api_key_here" \
-e MODEL="gemini-3.7-flash" \
-v ~/.kube/config:/app/.kube/config:ro \
talkopsai/k8s-autopilot:latest
The A2A/HTTP server will be accessible at http://localhost:10102.
You do not need to restart containers or edit files whenever you make changes. All major settings can be updated directly from the Settings UI at runtime:
| Variable | Description | Default / Options |
|---|---|---|
GOOGLE_API_KEY | API key for Google Gemini models | — |
OPENAI_API_KEY | API key for OpenAI models | — |
ANTHROPIC_API_KEY | API key for Anthropic Claude models | — |
MODEL | Active model (provider:model or bare name) | gemini-3.7-flash |
MODEL_PROVIDER | Provider override (auto-detected if omitted) | google_genai, openai, anthropic, ollama, etc. |
REASONING_EFFORT | Extended reasoning budget | low, medium, high, max (Default: medium) |
APPROVAL_MODE | Human approval safety mode | manual, auto, yolo (Default: manual) |
KUBECONFIG | Path to the mounted cluster kubeconfig | /app/.kube/config |
CHECKPOINT_BACKEND | Storage backend for sessions and configuration | auto, sqlite, postgres (Default: auto) |
POSTGRES_URI | PostgreSQL connection string | postgresql://user:pass@host:5432/db |
LOG_LEVEL | Application logging level | INFO, DEBUG, WARNING, ERROR |
Supports more than 20 model providers out of the box: Google GenAI, Anthropic, OpenAI, DeepSeek, Groq, Ollama (offline local models), Azure OpenAI, Google Vertex AI, AWS Bedrock, OpenRouter, Mistral AI, Together AI, Fireworks AI, and LiteLLM.
| Variable | Default Endpoint | System |
|---|---|---|
PROMETHEUS_BASE_URL | http://prometheus-operated.monitoring.svc:9090 | Prometheus Server |
ALERTMANAGER_BASE_URL | http://alertmanager-operated.monitoring.svc:9093 | Alertmanager Server |
LOKI_URL | http://localhost:3100 | Grafana Loki |
TEMPO_BASE_URL | http://localhost:3200 | Grafana Tempo |
ARGOCD_SERVER_URL | https://argocd-server.argocd.svc:443 | ArgoCD API Server |
TRAEFIK_API_URL | — | Traefik API / Dashboard |
Do I need all 11 MCP servers running in my cluster to use k8s-autopilot?
No. The 11 MCP servers are built into this image and start automatically as lightweight local subprocesses only when called. If a particular service (like Tempo or ArgoCD) is not installed in your cluster, the agent detects this and informs you that the specific capability is inactive. The rest of the framework continues to function normally.
Can I run k8s-autopilot completely offline?
Yes. You can run without external API keys by connecting k8s-autopilot to a local Ollama instance. Set MODEL=llama3.1 and MODEL_PROVIDER=ollama in your environment or select Ollama from the Settings UI.
How does this differ from Copilot, ChatGPT, or Claude Code?
General coding assistants generate code in an editor without direct awareness of your running cluster. k8s-autopilot is an operations runtime: it queries live telemetry, reads real pod health, formulates actionable operational plans, and safely executes them against your cluster with explicit human approval gates and post-execution verification.
What Kubernetes distributions can I manage?
Any CNCF-conformant Kubernetes cluster accessible from your kubeconfig. This includes cloud distributions (AWS EKS, Google GKE, Azure AKS), enterprise platforms (OpenShift), and local development clusters (Kind, Minikube, k3s, Talos).
Content type
Image
Digest
sha256:cbdb7f888…
Size
525 MB
Last updated
21 days ago
docker pull talkopsai/k8s-autopilot