Sign inSign up

talkopsai/k8s-autopilot

By talkopsai

•Updated 21 days ago

Multi-agent system for Kubernetes orchestration, progressive delivery, and live debugging.

Image
Machine learning & AI
Developer tools
Monitoring & observability
0

478

talkopsai/k8s-autopilot repository overview

k8s-autopilot Banner

An open-source, extensible AI operations framework for Kubernetes — with built-in operators for Helm, GitOps, progressive delivery, and full-stack observability.

Docker Pulls GitHub Repository Slack Community Discord License Python 3.12+


⁠Why k8s-autopilot?

Managing Kubernetes in production usually means jumping between a half-dozen browser tabs and terminal windows: pulling container logs, reading events, inspecting Prometheus dashboards, checking ArgoCD sync status, verifying Traefik routes, and looking up Helm revision history.

Most AI tools for DevOps try to solve this by acting as generic chatbot wrappers around kubectl. They lack real context on your cluster, cannot verify whether their proposed commands actually fixed anything, and running them with write permissions feels like an unnecessary risk to production stability.

k8s-autopilot takes a fundamentally different approach. It is an open-source, extensible operations framework that connects an autonomous multi-agent runtime directly to your live infrastructure through the Model Context Protocol (MCP). Instead of generating unverified scripts, it investigates your cluster, drafts a clear step-by-step plan, asks for your confirmation before changing anything, executes the work, and independently verifies the result against concrete acceptance criteria.


⁠The Framework Architecture

k8s-autopilot is not a single hardcoded assistant. It is a modular framework built on LangGraph⁠ and the Deep Agents SDK⁠. The core platform provides the orchestration layer, safety engine, verification loops, dynamic tool routing, and persistent memory.

                          User (Web UI / Chat / API)
                                      │
                                      ▼
                        Central Supervisor Agent
                                      │
             ┌────────────────────────┼────────────────────────┐
             ▼                        ▼                        ▼
     Built-in Operators        Marketplace Plugins       Custom Sub-agents
   (Helm, App, K8s, Obs)      (Claude, Codex, OpsCode)   (Dynamic Specialist Agents)
             │                        │                        │
             └────────────────────────┼────────────────────────┘
                                      │
                                      ▼
                       Model Context Protocol (MCP)
                  (11 built-in stdio servers + external HTTP/SSE)
                                      │
                                      ▼
                 Live Kubernetes Clusters & Infrastructure
⁠Core Framework Capabilities
⁠1. Extensible via Plugins and Marketplaces

The framework is designed so you decide what the agent can do. You can connect public or internal Git marketplaces and install plugins with a single click from the UI:

  • Broad Ecosystem Compatibility: Natively loads plugins built for Anthropic Claude, OpenAI Codex, OpsCode, or native k8s-autopilot manifests.
  • Vertical Plugins: Add cross-cutting skills directly into the central supervisor (e.g. security audits, cloud cost optimization, compliance checks).
  • Agent Plugins: Register entirely new, isolated specialist sub-agents that come bundled with their own dedicated skills and private MCP tool connections.
⁠2. Universal Tool Connectivity via MCP

Every integration in k8s-autopilot is built on the open Model Context Protocol (MCP)⁠.

  • Connect any MCP-compliant tool server via stdio, HTTP, or SSE directly from the Settings UI.
  • All 11 built-in infrastructure tools (Kubernetes, Helm, ArgoCD, Argo Rollouts, Traefik, Prometheus, Alertmanager, Loki, Tempo, OpenTelemetry, GitHub) communicate over lightweight, in-process stdio connections.
  • Connections follow a Just-In-Time (JIT) lifecycle: servers spin up only when an operator needs them and terminate once the task is complete, keeping resource usage minimal.
⁠3. Goal Tracking and Automated Rubric Grading

Most agents consider a task "done" as soon as an API call returns a 200 OK. k8s-autopilot treats operational tasks as verifiable goals:

  • Acceptance Criteria Generation: When given a non-trivial objective, the agent first formulates a concrete, testable rubric (e.g., "StatefulSet has 3/3 healthy replicas", "PVCs are bound", "ServiceMonitor is actively scraped by Prometheus").
  • Tactical Progress Checklist: Maintains an active checklist in the UI so you can follow real-time execution steps.
  • Independent Rubric Grading: When the worker agent finishes, an independent grader model inspects live cluster state against the criteria.
  • Self-Correction Loop: If any criterion fails, the grader generates a deficiency report, and the agent iterates until the work genuinely passes verification.
⁠4. Human-in-the-Loop (HITL) Safety Engine

Safety is the core design priority of the framework. You control how much autonomy the agent has:

  • Manual Mode (Default): The agent pauses and renders an interactive approval card before any mutating operation. Nothing touches your cluster until you click Approve.
  • Auto Mode: Read-only operations run instantly. Dangerous operations (deletions, namespace drains, production mutations) always pause for approval. Borderline operations are evaluated by an AI safety classifier with human fallback.
  • YOLO Mode: Unrestricted execution designed strictly for throwaway local environments or CI pipelines.
⁠5. Multi-Layer Defense-in-Depth
  • Semantic 4-Tier Profiling: Every tool across all MCP servers is profiled at startup into Tier 1 (Read-only), Tier 2 (Low-impact), Tier 3 (Mutating), or Tier 4 (Destructive).
  • Shell AST Scanner: Inspects piped commands and chained subshells to catch dangerous operations hidden inside seemingly benign command lines.
  • Invisible Unicode Detection: Scans and strips zero-width spaces, homoglyphs, and bidirectional override characters that could mislead human reviewers.
  • SSRF and Network Guard: Rejects requests to internal IP ranges, private subnets, and cloud metadata endpoints (169.254.169.254).
  • Headless Guard: When running unattended (via Slack or CI/CD), all mutating actions without explicit read-only annotations are automatically blocked.
⁠6. Persistent Memory and Operational Context

The agent remembers what it learns about your environments across conversations using AGENTS.md:

  • Stores cluster naming standards, preferred namespaces, ingress classes, and team conventions.
  • Preserves context across thread restarts so you never have to repeat your environment layout.
  • Protects critical configuration blocks from accidental agent modification.
⁠7. Interactive UI (A2A Protocol & A2UI)

The web interface pairs the open Agent-to-Agent (A2A) communication protocol with Agent-to-User Interface (A2UI) visual blueprints:

  • Interactive Approval Cards: Displays proposed changes, risk level badges, parameter summaries, and justification before execution.
  • Live Command Execution Drawers: Expandable execution logs with timers and raw stdout/stderr output.
  • Observability Visualizations: Renders interactive metric line charts, grouped log tables, and distributed trace waterfall timelines directly inside the conversation stream.

⁠Built-in Operators

Out of the box, k8s-autopilot ships with four domain operators that handle the most common Kubernetes workflows:

OperatorFocus AreaPrimary Capabilities
Kubernetes OperatorCore Cluster OperationsDiagnoses CrashLoopBackOff and OOMKilled pods; inspects container logs; safe pod exec and ephemeral debugging; scales deployments; audits RBAC roles and service accounts; switches multi-cluster contexts.
Helm OperatorPackage LifecycleDiscovers, installs, upgrades, and uninstalls releases; validates values against JSON schemas; dry-run preview; preserves values on upgrade; uses verified revision history for rollbacks; generates Helm charts and commits them to Git.
App OperatorGitOps & Progressive DeliveryManages ArgoCD application lifecycle and sync debugging; converts Deployments to Argo Rollouts with zero downtime; executes canary and blue-green deployments with automated Prometheus metric analysis; configures Traefik edge routing, weighted splitting, and middlewares.
Observability OperatorFull-Stack MonitoringTranslates natural language into PromQL, LogQL, and TraceQL queries; cross-pillar incident correlation across Prometheus, Alertmanager, Loki, and Tempo; manages exporter lifecycles and OpenTelemetry auto-instrumentation pipelines; triage alerts and silences with blast-radius preview.

⁠Quick Start

The automated script checks prerequisites (Docker, Docker Compose, kubeconfig), sets up ~/.k8s-autopilot/, pulls the images, and starts the stack:

curl -LsSf https://raw.githubusercontent.com/talkops-ai/k8s-autopilot/main/scripts/install.sh | bash

Once started, open http://localhost:8888⁠ in your browser.


⁠Option 2: Docker Compose

You can launch the complete setup — including the PostgreSQL persistence backend and the TalkOps Web UI — using Docker Compose:

1. Create a docker-compose.yml file:

services:
  postgres:
    image: postgres:16-alpine
    container_name: k8s-autopilot-postgres
    restart: unless-stopped
    environment:
      - POSTGRES_USER=k8s_autopilot
      - POSTGRES_PASSWORD=${POSTGRES_PASSWORD:-TalkOpsAutopilot2026SecureDB}
      - POSTGRES_DB=k8s_autopilot
    ports:
      - "5432:5432"
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U k8s_autopilot -d k8s_autopilot"]
      interval: 5s
      timeout: 5s
      retries: 5
    networks:
      - k8s-autopilot-net

  k8s-autopilot:
    image: talkopsai/k8s-autopilot:latest
    container_name: k8s-autopilot
    ports:
      - "10102:10102"
    depends_on:
      postgres:
        condition: service_healthy
    environment:
      - POSTGRES_URI=postgresql://k8s_autopilot:${POSTGRES_PASSWORD:-TalkOpsAutopilot2026SecureDB}@postgres:5432/k8s_autopilot?sslmode=disable
      - KUBECONFIG=/app/.kube/config
      # Optional: set an API key here, or configure it via the Settings UI later
      - GOOGLE_API_KEY=${GOOGLE_API_KEY}
      - MODEL=${MODEL:-gemini-3.7-flash}
      - REASONING_EFFORT=${REASONING_EFFORT:-medium}
    volumes:
      # Mount your kubeconfig so the agent can interact with your cluster
      - ${HOME}/.kube/config:/app/.kube/config:ro
      # Optional: mount local workspace for generated charts
      - ./workspace/helm-charts:/app/workspace/helm-charts
    restart: unless-stopped
    networks:
      - k8s-autopilot-net

  talkops-ui:
    image: talkopsai/talkops:latest
    container_name: talkops-ui
    environment:
      - K8S_AGENT_URL=http://localhost:10102
      - TALKOPS_ENABLE_LOGGING=false
      - VITE_ENVIRONMENT=production
    ports:
      - "8888:80"
    depends_on:
      - k8s-autopilot
    restart: unless-stopped
    networks:
      - k8s-autopilot-net

networks:
  k8s-autopilot-net:
    driver: bridge

volumes:
  postgres_data:

2. Start the stack:

docker compose up -d

3. Connect and configure:

  1. Open http://localhost:8888⁠ in your browser.
  2. Click Quick Connect (http://localhost:10102).
  3. Click Setup on the agent card to open the Settings panel.
  4. Go to Auth & Keys and add your API key (Google Gemini, Anthropic, OpenAI, etc.).
  5. Click Launch and start interacting with your cluster.

⁠Option 3: Standalone Container (docker run)

To run the agent API directly with your local kubeconfig and an API key:

docker run -d \
  --name k8s-autopilot \
  -p 10102:10102 \
  -e GOOGLE_API_KEY="your_api_key_here" \
  -e MODEL="gemini-3.7-flash" \
  -v ~/.kube/config:/app/.kube/config:ro \
  talkopsai/k8s-autopilot:latest

The A2A/HTTP server will be accessible at http://localhost:10102.


⁠Configuration and Runtime Management

You do not need to restart containers or edit files whenever you make changes. All major settings can be updated directly from the Settings UI at runtime:

  • Auth & Keys: Add, rotate, or remove API keys across 20+ supported LLM providers.
  • Model Switching: Switch models or tweak reasoning effort levels mid-conversation.
  • Backend Storage: Hot-swap between SQLite (zero-config local storage) and PostgreSQL (production multi-replica setups) with an instant connectivity test.
  • MCP Servers: View active tool servers, toggle servers on/off, or add external servers over stdio, HTTP, or SSE.
  • Plugins: Browse connected marketplaces, install new skills, and manage active sub-agents.
  • Cluster Connectivity: Update kubeconfig paths, target contexts, default namespaces, and sandbox execution environments.
⁠Core Environment Variables
VariableDescriptionDefault / Options
GOOGLE_API_KEYAPI key for Google Gemini models—
OPENAI_API_KEYAPI key for OpenAI models—
ANTHROPIC_API_KEYAPI key for Anthropic Claude models—
MODELActive model (provider:model or bare name)gemini-3.7-flash
MODEL_PROVIDERProvider override (auto-detected if omitted)google_genai, openai, anthropic, ollama, etc.
REASONING_EFFORTExtended reasoning budgetlow, medium, high, max (Default: medium)
APPROVAL_MODEHuman approval safety modemanual, auto, yolo (Default: manual)
KUBECONFIGPath to the mounted cluster kubeconfig/app/.kube/config
CHECKPOINT_BACKENDStorage backend for sessions and configurationauto, sqlite, postgres (Default: auto)
POSTGRES_URIPostgreSQL connection stringpostgresql://user:pass@host:5432/db
LOG_LEVELApplication logging levelINFO, DEBUG, WARNING, ERROR
⁠Supported Model Providers

Supports more than 20 model providers out of the box: Google GenAI, Anthropic, OpenAI, DeepSeek, Groq, Ollama (offline local models), Azure OpenAI, Google Vertex AI, AWS Bedrock, OpenRouter, Mistral AI, Together AI, Fireworks AI, and LiteLLM.

⁠Integration Endpoints (Optional Overrides)
VariableDefault EndpointSystem
PROMETHEUS_BASE_URLhttp://prometheus-operated.monitoring.svc:9090Prometheus Server
ALERTMANAGER_BASE_URLhttp://alertmanager-operated.monitoring.svc:9093Alertmanager Server
LOKI_URLhttp://localhost:3100Grafana Loki
TEMPO_BASE_URLhttp://localhost:3200Grafana Tempo
ARGOCD_SERVER_URLhttps://argocd-server.argocd.svc:443ArgoCD API Server
TRAEFIK_API_URL—Traefik API / Dashboard

⁠Frequently Asked Questions

Do I need all 11 MCP servers running in my cluster to use k8s-autopilot?
No. The 11 MCP servers are built into this image and start automatically as lightweight local subprocesses only when called. If a particular service (like Tempo or ArgoCD) is not installed in your cluster, the agent detects this and informs you that the specific capability is inactive. The rest of the framework continues to function normally.

Can I run k8s-autopilot completely offline?
Yes. You can run without external API keys by connecting k8s-autopilot to a local Ollama instance. Set MODEL=llama3.1 and MODEL_PROVIDER=ollama in your environment or select Ollama from the Settings UI.

How does this differ from Copilot, ChatGPT, or Claude Code?
General coding assistants generate code in an editor without direct awareness of your running cluster. k8s-autopilot is an operations runtime: it queries live telemetry, reads real pod health, formulates actionable operational plans, and safely executes them against your cluster with explicit human approval gates and post-execution verification.

What Kubernetes distributions can I manage?
Any CNCF-conformant Kubernetes cluster accessible from your kubeconfig. This includes cloud distributions (AWS EKS, Google GKE, Azure AKS), enterprise platforms (OpenShift), and local development clusters (Kind, Minikube, k3s, Talos).


⁠Community and Documentation

Tag summary

Content type

Image

Digest

sha256:cbdb7f888…

Size

525 MB

Last updated

21 days ago

docker pull talkopsai/k8s-autopilot