Sign inSign up

psyb0t/aicodebox

By psyb0t

•Updated about 4 hours ago

Base Docker image that wraps any terminal AI coding agent in an HTTP API, an OpenAI-compatible en...

Image
0

10K+

psyb0t/aicodebox repository overview

Source⁠ | Project page⁠

⁠docker-aicodebox

CI version license Docker Pulls

The agent-agnostic foundation for putting any terminal-shaped AI coding agent on the network. Pick your poison — claude-code, pi, opencode, hermes, whatever vibes — bolt on a 20-line adapter, and out the other end you get an HTTP API, an OpenAI-compatible chat completions endpoint, an MCP server, a Telegram bot, and a cron scheduler that fires the agent on whatever schedule you can dream up. One container. Same surfaces. Swap the brain.

You don't fork this. You FROM it.

FROM psyb0t/aicodebox

RUN npm install -g @earendil-works/[email protected]
COPY mypkg /opt/mypkg
RUN pip3 install --break-system-packages /opt/mypkg

ENV AICODEBOX_ADAPTER=mypkg.adapter:MyAdapter \
    AICODEBOX_AGENT_BINARY=pi

That's it. The base owns the surfaces. Your adapter translates "run this prompt" into whatever your agent's CLI expects. New agent lands in an afternoon.

aicodebox is a base image, not a host launcher. Use a child wrapper such as claudebox⁠, codexbox⁠, or pibox⁠ for local agent work. Those wrappers mount the requested workspace and the matching agent state. Child agents call a sibling wrapper directly when it is installed beside their own wrapper. They must not reconstruct its Docker command or mount another agent's state themselves.

⁠Table of Contents

⁠What's in the box

LayerThe goods
OSUbuntu 24.04. aicode user (UID 1000), passwordless sudo, docker group.
RuntimesNode.js 24.20.0 LTS (for agents that ship as npm), Python 3.14.7, Docker CE + buildx + compose (in case your agent needs to spawn containers).
Packageaicodebox — the adapter contract + four mode dispatchers (api / telegram / cron / mcp). Pure Python, zero side effects until you boot a mode.
ModesAll optional, all opt-in via env vars. Run none or one per container. Exception: telegram + cron can share a container — cron runs in-thread inside the telegram process.
AuthAICODEBOX_API_MODE_TOKEN gates API mode; AICODEBOX_MCP_MODE_TOKEN gates MCP. Single bearer per surface, no fallback between them. Empty = no auth. Telegram has its own allowlist.
StatePer-chat overrides + cron history go under $HOME/.aicodebox/, along with your own bin/ and init.d/ (see Init scripts and user bin⁠). Bind-mount that path if you want it to outlive the container. The package itself stores nothing.

psyb0t/aicodebox:latest is deliberately small. The matching psyb0t/aicodebox:latest-full variant adds the shared development toolchain for agent images that need it: Go and its language servers and linters, pinned Node and Python developer tools, editors and diagnostics, database clients, and gh, Terraform, kubectl, and Helm. It is built after the minimal image in CI, so the full tag inherits the minimal image from the same release.

⁠The adapter contract

Everything routes through one interface. You implement it once per agent.

# mypkg/adapter.py
from aicodebox.adapters.base import AgentAdapter, RunRequest, RunResult, StreamEvent

class MyAdapter(AgentAdapter):
    name = "my-agent"
    available_models = ["fast", "smart"]
    available_thinking_levels = ["off", "low", "high"]

    def build_argv(self, req: RunRequest) -> list[str]:
        argv = ["my-agent", "-p", req.prompt]
        if req.model: argv += ["--model", req.model]
        if req.workspace: argv += ["--cwd", req.workspace]
        return argv

    def parse_output(self, stdout: str, req: RunRequest) -> RunResult:
        # Required. Convert raw stdout into a normalized result. raw_stdout /
        # raw_stderr / exit_code are filled in by the runner — set ``text``
        # (and optionally ``parsed`` / ``usage`` / ``session_id``).
        return RunResult(text=stdout.strip(), raw_stdout="", raw_stderr="", exit_code=0)

    # Optional: structured stream events. The default is "one delta per
    # stdout line" — override only if your binary emits a JSON event stream
    # and you want per-token / per-tool granularity in OAI streaming.
    def parse_stream_event(self, line: str, req: RunRequest) -> StreamEvent | None:
        return StreamEvent(type="delta", text=line + "\n") if line else None

Full event retention is runner-owned. With eventMode: "full", it keeps each non-empty stdout line. JSON objects remain provider records. Diagnostics, malformed JSON, and JSON values that are not objects remain raw line records.

ENV AICODEBOX_ADAPTER=mypkg.adapter:MyAdapter

The package gets resolved at first call, cached for the process lifetime. Every mode pulls the same adapter — what gets exposed over HTTP / MCP / Telegram / cron is exactly what your build_argv knows how to drive.

⁠Modes

Modes are controlled by env vars. Set the flag, the entrypoint starts that mode. No flag, no mode. Foreground modes (API / Telegram / Cron) are mutually exclusive — except telegram + cron, which share a process (cron runs in-thread inside telegram). API wins if set alongside anything else. MCP mode is independent — it coexists with any foreground mode, served on its own port (or mounted at /mcp/ inside API).

⁠API mode

AICODEBOX_API_MODE=1. Boots a FastAPI server on :8080 (override with AICODEBOX_API_MODE_PORT) with:

Required: AICODEBOX_AVAILABLE_MODELS=<csv> — /openai/v1/models needs a real list, and there's no safe fallback (the adapter name isn't a model name). API mode refuses to boot without it. Pick the model ids your configured provider actually serves.

  • POST /run — sync agent run. jsonSchema and event collection are independent:

    • No jsonSchema, eventMode: "none" returns {runId, workspace, exitCode, text}.
    • "jsonSchema": {...} validates the final text and adds json, sessionId, usage, and attempts. Failed validation retries up to three times and returns parseError plus jsonRetries if every attempt fails.
    • "eventMode": "full" adds the complete native agent event stream as events, including reasoning, tool lifecycle, usage, and provider metadata where the CLI exposes them. Each record has the stable envelope {sequence, attempt, backend, eventType, event}. event is the untouched provider record.
    • "eventMode": "none" suppresses events, including for a schema request. "eventMode": "auto" is the default and keeps compatibility by enabling events for a schema request or "outputFormat": "json-verbose".

    "outputFormat": "text" and "outputFormat": "json" remain accepted legacy inputs but do not change the /run response shape. New callers should select event retention with eventMode; only json-verbose has a compatibility effect while eventMode is auto.

    usage is the sum across every schema retry. attempts is the per-attempt breakdown [{index, usage, exitCode, parseError}]. Set includeRaw to also receive stdout and stderr; stderr is always included when exitCode != 0.

    {
      "prompt": "Inspect the workspace and return a short report.",
      "eventMode": "full",
      "jsonSchema": {"type": "object"}
    }
    
  • POST /run with "async": true or "fireAndForget": true returns {runId, workspace, status: "running", fireAndForget} immediately

  • GET /run/result?runId=<id> polls an async run. A completed response includes the synchronous result fields plus status.

  • DELETE /run/{id} — kill an in-flight run

  • GET|PUT|DELETE /files/{path} — workspace file CRUD

  • POST /openai/v1/chat/completions — OpenAI-compatible (streaming + non-streaming). Plug it into anything that speaks OpenAI. Schema-validated JSON output: stock OpenAI clients drive it via the standard response_format body field ({"type":"json_object"} for permissive, {"type":"json_schema","json_schema":{"name":"...","schema":{...}}} for structured outputs); the proprietary x-aicodebox-json-schema header is supported as a fallback. Body field wins if both are set. Schema mode runs up to 3 self-correction retries — success → canonical JSON in message.content, retries exhausted → 422, agent process crash → 500 with exit code + stderr in detail, combined with stream=true → buffered SSE (see below). Additional RunSpec knobs via x-aicodebox-* headers: workspace, continue, append-system-prompt, resume, extra-args, timeout-seconds, tools-allowlist, no-tools. Malformed header values surface as 400 with the offending header name. When schema mode runs retries, usage is the sum across all attempts (input/output/total/cache fields all summed), and the envelope carries a vendor-extension aicodebox_attempts: [{index, usage, exitCode, parseError}, ...] so callers can see the per-attempt breakdown (OAI-only clients ignore the unknown field). Cheap retries via session continuation: schema requests that omit x-aicodebox-workspace get a per-request ephemeral workspace under /tmp/aicodebox/<uuid>/ (cleaned up in finally); retries then run with no_continue=False and a minimal corrective prompt (error + directive + schema) instead of replaying the full original input, cutting per-retry input cost roughly 100x on large prompts. Callers that DO provide their own workspace fall back to fresh-session retries that re-state the original task — safe across any workspace, but more expensive. Client-executed tool calling: send the standard tools array (+ optional tool_choice) and the machine acts as a plain function-calling model — it responds with tool_calls + finish_reason:"tool_calls" when it wants a tool, your client runs the tool and sends the role:"tool" result back, and the loop continues (stateless, resend full history each round, exactly like OpenAI). tool_choice supports auto/none/required/{type:"function",function:{name}}. tools + response_format compose (agentic-then-structured): a tool-call turn returns tool_calls/finish_reason:"tool_calls" and is not schema-checked, while the model's final answer turn is validated against the schema (with retry) and returned as canonical JSON — so a multi-tool flow can end in a structured reply. Streaming for tool/schema modes (stream=true with tools or response_format) is served as buffered SSE: the full answer is computed, then replayed as a single-shot text/event-stream (opening role chunk → one content/tool_calls delta with the required index → finish chunk → data: [DONE]) — a valid stream, just not token-incremental. Plain chat still streams incrementally. In tool mode the harness's own internal tools default off (pure function-caller); send x-aicodebox-no-tools: 0 to re-enable the hybrid (internal + client tools together).

    For a stream that also needs every native provider record, send "stream_options": {"include_aicodebox_events": true}. The response adds named aicodebox.native SSE records with {sequence, attempt, backend, eventType, event} before the ordinary OpenAI chunks. Normal content chunks and [DONE] stay unchanged. The option requires stream: true.

  • GET /openai/v1/models — model list from the adapter

  • POST /mcp/ — MCP server (mounted only when AICODEBOX_MCP_MODE=1; auth via AICODEBOX_MCP_MODE_TOKEN, separate from the API bearer)

Bearer auth for the API surface: AICODEBOX_API_MODE_TOKEN=<one-token>. Single token, no rotation list. Empty = no auth.

⁠Telegram mode

AICODEBOX_TELEGRAM_MODE=1 + AICODEBOX_TELEGRAM_MODE_TOKEN=<bot:token>. Drop the bot into a chat, talk to it, get answers. Features:

  • Text in → agent run → response chunked + Markdown→HTML rendered for Telegram.
  • File uploads (document / photo / video / voice) land in the chat's workspace.
  • [SEND_FILE: relative/path] in agent output delivers workspace files back as Telegram attachments.
  • Per-chat overrides: /model, /effort, /system_prompt, /append_system_prompt. Persisted to disk.
  • /cancel kills the in-flight run for the chat. /reload re-reads the yaml. /config dumps merged chat config. /fetch <path> downloads a workspace file. /status lists busy chats.
  • Replies to cron-fired messages inject the job's instruction + result so follow-ups make sense.

Config lives at $HOME/.aicodebox/telegram.yml:

allowed_chats: [-100123, 42]
default:
  model: glm-4.5-air
  workspace: shared
chats:
  -100123:
    workspace: alpha
    model: claude-sonnet
    allowed_users: [10, 20]
⁠Cron mode

AICODEBOX_CRON_MODE=1 + AICODEBOX_CRON_MODE_FILE=/path/to/cron.yaml. It accepts five field schedules for minute resolution and six field schedules for second resolution. Jobs can select safe relative workspace directories and optionally notify Telegram.

jobs:
  - name: morning-report
    schedule: "0 0 9 * * *"
    instruction: |
      Summarize yesterday's git activity in {workspace}.
    workspace: shared
    telegram_chat_id: -100123
    model: claude-sonnet

Each run gets its own history directory under $HOME/.aicodebox/cron/history/<workspace-slug>/<YYYYMMDD-HHMMSS>-<job>/ with meta.json, stdout.log, stderr.log, result.txt, and telegram.json when notified. The scheduler also appends a summary to $HOME/.aicodebox/cron/<job>.jsonl. Set AICODEBOX_CRON_MODE_HISTORY_DIR to move the whole cron state root, including the Telegram reply metadata.

⁠MCP mode

AICODEBOX_MCP_MODE=1. Exposes the MCP (Model Context Protocol) surface. Coexists with any foreground mode:

ForegroundMCP placement
API mode (AICODEBOX_API_MODE=1)mounted at /mcp/ on the API port — no extra process
Telegram / Cron / passthroughruns as a sidecar uvicorn at / on AICODEBOX_MCP_MODE_PORT (default 8081)

Auth: AICODEBOX_MCP_MODE_TOKEN=<one-token> — bearer token in the Authorization: Bearer … header, or ?apiToken=… for clients that can't set headers. Empty = no auth. No fallback to API_MODE_TOKEN — MCP is its own surface with its own bearer.

Point Claude Desktop / Cursor / whatever at http://host:8080/mcp/ in API mode or http://host:8081/ in standalone mode. The agent shows up as a set of tools (run_prompt, list_files, read_file, write_file, delete_file).

MCP keeps DNS rebinding protection enabled. Loopback hosts and origins work by default. A reverse proxy, tunnel, or public DNS name must set AICODEBOX_MCP_MODE_ALLOWED_HOSTS to the exact Host values it forwards and AICODEBOX_MCP_MODE_ALLOWED_ORIGINS to exact browser origins, including their schemes. Do not disable this protection or use broad wildcards for an internet-facing endpoint.

⁠Configuration

Everything's an env var. The base sets sane defaults, your child image overrides.

Env var convention: <MODE>_MODE is the on/off flag for that mode; <MODE>_MODE_<KNOB> is its config. Vars that aren't mode-scoped (workspace, adapter, container) are bare.

⁠Adapter & container
VarDefaultWhat it does
AICODEBOX_ADAPTERrequiredpkg.module:Class reference to your AgentAdapter subclass
AICODEBOX_AGENT_BINARYrequiredName of the agent's CLI binary (for which checks, version reports)
AICODEBOX_WORKSPACE/workspaceRoot dir for all per-chat / per-job workspaces
AICODEBOX_CONTAINER_NAMEaicodeboxDisplay name in /status, logs, and per-container state files
AICODEBOX_AVAILABLE_MODELS—Required for API mode. CSV list returned by /openai/v1/models and shown in the telegram /model picker. API mode refuses to boot without it; telegram /model picker degrades to a "set this env var" reply.
AICODEBOX_AVAILABLE_EFFORTSadapter listOverride the effort/--thinking list exposed via /effort (comma-separated)
⁠Mode flags
VarDefaultWhat it does
AICODEBOX_API_MODE0Boot the HTTP API server (foreground)
AICODEBOX_TELEGRAM_MODE0Boot the Telegram bot (foreground)
AICODEBOX_CRON_MODE0Boot the cron scheduler (foreground; runs in-thread if telegram is also on)
AICODEBOX_MCP_MODE0Expose the MCP server — mounted at /mcp/ in API mode, or as a sidecar elsewhere
⁠API mode config
VarDefaultWhat it does
AICODEBOX_API_MODE_PORT8080Port the API server binds to
AICODEBOX_API_MODE_TOKENemptyBearer token for the API surface. Empty = no auth
⁠Telegram mode config
VarDefaultWhat it does
AICODEBOX_TELEGRAM_MODE_TOKEN—Bot token from @BotFather
AICODEBOX_TELEGRAM_MODE_CONFIG$HOME/.aicodebox/telegram.ymlPath to the telegram config yaml
AICODEBOX_TELEGRAM_MODE_OVERRIDES$HOME/.aicodebox/telegram_overrides.jsonPer-chat override store (model/effort/system prompts)
⁠Cron mode config
VarDefaultWhat it does
AICODEBOX_CRON_MODE_FILE—Path to the cron yaml
AICODEBOX_CRON_MODE_HISTORY_DIR$HOME/.aicodebox/cronCron state root for run artifacts, job summaries, and Telegram reply metadata
⁠MCP mode config
VarDefaultWhat it does
AICODEBOX_MCP_MODE_PORT8081Port the sidecar MCP server binds to (ignored when MCP is mounted inside API)
AICODEBOX_MCP_MODE_TOKENemptyBearer token for MCP. Empty = no auth. No fallback to API_MODE_TOKEN
AICODEBOX_MCP_MODE_ALLOWED_HOSTSloopback hostsComma-separated Host values accepted by MCP. Add each reverse-proxy host name.
AICODEBOX_MCP_MODE_ALLOWED_ORIGINSloopback HTTP originsComma-separated browser origins accepted by MCP. Add each proxy origin that sends browser requests.

⁠Init scripts and user bin

The entrypoint runs init scripts once per container, the first time it starts. Restarting the container does not run them again; a new container does, even when it mounts the same $HOME/.aicodebox. The record of the run lives in the container filesystem at /var/lib/aicodebox/init-done.

  1. /aicodebox-init.d/*.sh, which the child image bakes in, run first.
  2. $HOME/.aicodebox/init.d/*.sh, your own, run next. Bind-mount $HOME/.aicodebox to supply them.

Both sets run in filename order as aicode, which has passwordless sudo, with $HOME/.aicodebox/bin already on PATH. A script that exits non-zero is logged as [entrypoint] init script <path> failed, and the rest still run.

Executables in $HOME/.aicodebox/bin are on PATH ahead of everything else, for the agent in every mode, for init scripts, and for docker exec shells.

⁠Child image recipe

Minimal adapter that wires up an npm-shipped agent:

FROM psyb0t/aicodebox:latest

# Your agent — pin the version.
ARG AGENT_VERSION=0.74.0
RUN npm install -g @your-org/your-agent@${AGENT_VERSION}

# Your adapter package — implements aicodebox.adapters.base.AgentAdapter.
COPY your_adapter /opt/your_adapter
RUN pip3 install --no-cache-dir --break-system-packages /opt/your_adapter

ENV AICODEBOX_ADAPTER=your_adapter.adapter:YourAdapter \
    AICODEBOX_AGENT_BINARY=your-agent

Boot it:

docker run --rm -p 8080:8080 \
  -e AICODEBOX_API_MODE=1 \
  -e AICODEBOX_API_MODE_TOKEN=$(openssl rand -hex 16) \
  -v "$PWD/workspace:/workspace" \
  your/child-image:latest

A reference child image lives at psyb0t/pibox⁠ — wraps pi-coding-agent⁠ and uses this base verbatim.

⁠Full image

Use the full variant only when the agent needs a general development toolchain:

FROM psyb0t/aicodebox:latest-full

It keeps the same entrypoint, workspace contract, Python package, and mode surfaces as latest. The difference is the toolchain. A child image should select one parent image per variant and install only its agent-specific CLI, adapter, entrypoint, and configuration.

⁠Development

make help            # list targets
make build           # docker build .
make build-full      # build aicodebox:latest-full on the matching local minimal tag
make build-all        # build both variants
make test            # python unit tests (199 cases, adapter contract, modes, helpers)
make test-unit       # same as test
make test-integration # build and verify the init.d and user bin/ hooks
make test-full-image # build full and verify every documented CLI tool
make lint            # flake8 + pyright
make format          # isort + black
make clean           # nuke caches + the built image

Tests run in-process — no docker required. The suite stubs out the adapter via AICODEBOX_ADAPTER=aicodebox.tests.conftest:_StubAdapter so the modes can be exercised without a real agent on disk.

For integration testing with a real agent + real Telegram chat, see the e2e harness in the pibox⁠ repo — it uses psyb0t/telethon-plus⁠ as a userbot driver.

⁠License

WTFPL — see LICENSE⁠. Do what the fuck you want.

Tag summary

Content type

Image

Digest

sha256:998bc14a7…

Size

450.7 MB

Last updated

about 5 hours ago

docker pull psyb0t/aicodebox