Sign inSign up

thebuildguild/cli-bridge

By thebuildguild

•Updated 30 days ago

CLI Bridge API - HTTP gateway for AI CLIs (Codex & Claude) and Whisper Audio Transcription

Image
API management
Machine learning & AI
Developer tools
1

5.4K

thebuildguild/cli-bridge repository overview

⁠CLIBridge

HTTP gateway that turns the OpenAI Codex CLI or Anthropic Claude Code CLI into a standard OpenAI-compatible API — now with a 100% local, self-hosted Whisper transcription endpoint too.

Use your own ChatGPT Plus or Claude Pro subscription. CLIBridge handles authentication, streaming, image inputs, audio transcription, rate limiting, health monitoring, and SaaS-based activation.

Audio sent to POST /v1/audio/transcriptions is transcribed entirely inside the container by a bundled whisper.cpp — it is never sent to OpenAI, Anthropic, or any other external service.

Source / issues: https://github.com/buildtheguild/cli-bridge⁠

⁠Tags

TagCLI backendsNotes
latest, 2.3.0Codex + ClaudeIncludes whisper.cpp + baked-in Whisper model
claude, 2.3.0-claudeClaude onlyIncludes whisper.cpp + baked-in Whisper model
codex, 2.3.0-codexCodex onlyIncludes whisper.cpp + baked-in Whisper model

As of 2.0.0, every tag also bundles whisper.cpp and a baked-in Whisper model (base by default), which adds roughly the size of that model (~150 MB for base) on top of prior image sizes. As of 2.1.0, other models (tiny/small/medium/large-v3-turbo/large-v3) can be downloaded and switched to entirely from the dashboard or API — no rebuild needed. See "Audio transcription" below.

Use a versioned tag in production.

Full image and BACKEND: The latest / full image has both CLIs pre-installed but still routes all traffic through a single backend selected at runtime via BACKEND=codex or BACKEND=claude. There is no simultaneous dual-backend routing — the full image is a build-time convenience so you can switch providers without rebuilding the image.

⁠Requirements

  • Docker
  • A valid CLIBridge subscription
  • A CLI account for the backend you want to use

⁠Quick start

⁠1. Create .env

Pick the backend you want to use.

Codex:

BRIDGE_TOKEN=your-long-random-secret-here
BACKEND=codex

Claude:

BRIDGE_TOKEN=your-long-random-secret-here
BACKEND=claude
⁠2. Create docker-compose.yml

If your .env was copied from a Windows host, it may contain a Windows binary path (e.g. CODEX_BIN=C:\Users\you\AppData\Roaming\npm\codex.cmd). The environment: block below overrides that back to the in-container binary name — environment: always wins over env_file for the same key, so the rest of .env is unaffected.

Codex:

services:
  cli-bridge:
    image: thebuildguild/cli-bridge:2.3.0-codex
    ports:
      - "3900:3900"
    env_file:
      - .env
    environment:
      PORT: 3900
      BACKEND: codex
      CODEX_BIN: codex
      API_DOCS_SERVERS: http://localhost:3900,https://your.domain.com
    volumes:
      - cli_data:/data
    restart: unless-stopped

volumes:
  cli_data:

Claude:

services:
  cli-bridge:
    image: thebuildguild/cli-bridge:2.3.0-claude
    ports:
      - "3900:3900"
    env_file:
      - .env
    environment:
      PORT: 3900
      BACKEND: claude
      CLAUDE_BIN: claude
      API_DOCS_SERVERS: http://localhost:3900,https://your.domain.com
    volumes:
      - cli_data:/data
    restart: unless-stopped

volumes:
  cli_data:

Combined (both CLIs installed, pick one backend at runtime):

services:
  cli-bridge:
    image: thebuildguild/cli-bridge:2.3.0
    ports:
      - "3900:3900"
    env_file:
      - .env
    environment:
      PORT: 3900
      BACKEND: codex  # or claude — switch anytime without rebuilding
      CODEX_BIN: codex
      CLAUDE_BIN: claude
      API_DOCS_SERVERS: http://localhost:3900,https://your.domain.com
    volumes:
      - cli_data:/data
    restart: unless-stopped

volumes:
  cli_data:
⁠3. Start
docker compose pull
docker compose up -d
⁠4. Activate

Open http://localhost:3900, click Activate, and approve the device in your browser.

⁠5. Authenticate (one-time)

Recommended — web wizard:

Open http://localhost:3900, go to the provider step, and click Authenticate. Complete the login in the browser tab that opens; the page detects the connection automatically — no terminal needed.

Alternative — CLI (headless/scripted setups):

Claude:

docker compose exec cli-bridge sh
claude auth login

Codex:

docker compose exec cli-bridge sh
codex login --device-auth
⁠6. Keeping the CLI itself up to date

The dashboard checks the npm registry for newer Codex/Claude CLI releases and shows an Update available banner with an Update button on the provider's card when one is found — no rebuild needed. Clicking it runs npm install -g <package>@latest inside the running container. It's manual on purpose: nothing updates itself.

This patches the running container's writable layer only. A normal docker compose up -d / docker compose restart keeps it; a docker compose up --force-recreate, an image pull, or a fresh deploy reverts to whatever CLI version is baked into the image. Rebuild and republish with a newer OPENAI_CODEX_VERSION/CLAUDE_CODE_VERSION build arg if you want the new version to stick across redeploys.

⁠7. Audio transcription (optional, plan-gated)

POST /v1/audio/transcriptions doesn't depend on your Codex/Claude subscription at all — model baked into the image, and nothing is sent to OpenAI, Anthropic, or any other external service. It does require your CLIBridge plan to include Whisper; if it doesn't, both this endpoint and GET /v1/audio/status return 403, and the dashboard hides the Whisper card entirely rather than showing it locked.

curl http://localhost:3900/v1/audio/transcriptions \
  -H "Authorization: Bearer $BRIDGE_TOKEN" \
  -F "[email protected]"

Check readiness anytime with GET /v1/audio/status (binary/model/ffmpeg health, live capacity, usage metrics — no auth-gated verbose health flag needed). If a transcription hangs or is taking too long, POST /v1/audio/cancel aborts whatever's currently running and frees the slot immediately instead of waiting out the timeout.

Multiple models, switchable without a rebuild: GET /v1/audio/models lists a catalog (tiny/base/small/medium/large-v3-turbo/large-v3) with size and downloaded/active state. POST /v1/audio/models/:name/select switches immediately if already downloaded, or downloads it from Hugging Face straight to your persistent volume and switches automatically once done — the dashboard's Whisper card exposes this as a dropdown with a live progress bar (and a Cancel button while downloading). DELETE /v1/audio/models/:name frees disk space. Downloaded models and the active selection survive container recreation and image updates.

Per-request override: pass model (e.g. model=large-v3-turbo) on POST /v1/audio/transcriptions to use a different already-downloaded model for just that one request, without changing the dashboard's default. An undownloaded or unknown model name returns a clear 400 rather than silently falling back or triggering a multi-minute download mid-request.

⁠Model selection after backend switches

As of 2.2.1, OpenAI-compatible chat requests are more forgiving when you've switched the bridge from one provider to the other but an upstream automation is still sending the old provider's model id.

Example: if the bridge is now running with BACKEND=codex but your n8n workflow still sends claude-sonnet-4-6, or the bridge is running with BACKEND=claude but the caller still sends gpt-5.5, the request can fall back to the active backend's default model instead of failing with Unknown model ....

This behavior is controlled by:

MODEL_FALLBACK_TO_DEFAULT_ON_UNKNOWN=true

It is enabled by default, including when the variable is missing from .env. Set it to false if you want strict model validation again.

When fallback happens, the successful JSON response includes a top-level model_fallback object so the caller can detect the mismatch:

{
  "model": "gpt-5.5",
  "model_fallback": {
    "requested": "claude-sonnet-4-6",
    "used": "gpt-5.5",
    "reason": "unknown_model",
    "message": "Requested model 'claude-sonnet-4-6' is not available for the active backend, so the server default model was used instead."
  }
}
⁠Emergency OpenRouter fallback

As of 2.2.2, the bridge can also fail over internally to OpenRouter for the main chat request path, but only if you opt in explicitly:

OPENROUTER_ENABLE_EMERGENCY_FALLBACK=true
OPENROUTER_API_KEY=...
OPENROUTER_DEFAULT_MODEL=openai/gpt-5

Optional:

OPENROUTER_FALLBACK_MODELS=anthropic/claude-sonnet-4.5,google/gemini-2.5-pro
OPENROUTER_TIMEOUT_MS=45000

This is intentionally emergency-only. It does not add any new public endpoint, and it does not activate for generic bridge errors. It only retries through OpenRouter when the active CLI request fails with a known hard condition such as:

  • missing CLI authentication
  • provider rate limiting (429)
  • provider quota exhaustion / usage-limit failure
  • provider timeout or unavailability

Successful responses include a provider_fallback object so the caller can see that OpenRouter handled the request:

{
  "provider_fallback": {
    "from": "codex",
    "to": "openrouter",
    "reason": "rate_limited",
    "model": "openai/gpt-5"
  }
}

This uses OpenRouter API credits, not your Codex/Claude subscription quota, so leave it disabled unless you intentionally want that emergency path.

Diagnosing failures: GET /v1/logs (and the dashboard's Logs tab) shows recent transcription/download failures with the actual error — e.g. the exact whisper-cli timed out after ... message — instead of just a bare failure count. It also covers Codex/Claude chat and streaming request failures, so a failed request shown in the Requests/Failure Rate metrics always has a matching explanation here. The periodic insights report (see INSIGHTS_* below) is excluded from those metrics — it's internal housekeeping, not user traffic — but its failures still land in GET /v1/logs, so a stale "not logged in" blip from before you authenticated won't leave a permanent-looking 100% failure rate on the dashboard.

⁠Configuration

Set these in .env (loaded via env_file) or directly under environment: in compose — environment: always wins on conflicts.

Required

  • BRIDGE_TOKEN — shared secret for all HTTP requests

Server

  • PORT — listen port (default 3000; examples here use 3900)
  • CORS_ORIGINS — comma-separated allowed origins (default: allow all)
  • ENABLE_API_DOCS — enable Swagger at /api-docs (default true)
  • API_DOCS_SERVERS — comma-separated server URLs shown in the Swagger UI server picker, e.g. http://localhost:3900,https://your.domain.com
  • API_DOCS_EXPANSION — Swagger expansion mode: full | list | none
  • APP_TITLE, APP_DESCRIPTION, APP_SERVICE_NAME — branding shown in Swagger/health
  • HEALTH_VERBOSE_ENABLED — expanded GET /v1/health?details=true response (default false)

Backend selection

  • BACKEND — codex (default) or claude. Independent of CODEX_BIN/CLAUDE_BIN: BACKEND picks which provider's request path is active, the *_BIN vars only say which binary that path spawns — the unused one is ignored.

Codex CLI (used when BACKEND=codex)

  • CODEX_BIN — binary name (default codex)
  • CODEX_TIMEOUT_MS, CODEX_DEFAULT_MODEL, CODEX_ALLOWED_MODELS, CODEX_SANDBOX, CODEX_ASK_FOR_APPROVAL, CODEX_ALLOW_SEARCH
  • CODEX_SKIP_GIT_REPO_CHECK — default true (the bridge runs as a service, not inside a project directory)
  • OPENAI_MODELS_HIDE_CODEX — default true

Claude CLI (used when BACKEND=claude)

  • CLAUDE_BIN — binary name (default claude)
  • CLAUDE_DEFAULT_MODEL (default claude-sonnet-4-6), CLAUDE_ALLOWED_MODELS

Emergency OpenRouter fallback (optional, disabled by default)

  • OPENROUTER_ENABLE_EMERGENCY_FALLBACK — enable internal emergency failover for chat requests
  • OPENROUTER_API_KEY — OpenRouter API key
  • OPENROUTER_DEFAULT_MODEL — model used for the emergency request
  • OPENROUTER_FALLBACK_MODELS — optional comma-separated OpenRouter fallback chain
  • OPENROUTER_TIMEOUT_MS — timeout for the emergency OpenRouter call (default 45000)
  • OPENROUTER_HTTP_REFERER, OPENROUTER_TITLE — optional attribution headers

Audio transcription (POST /v1/audio/transcriptions, always available — independent of BACKEND)

  • WHISPER_MODEL_PATH — path to the .bin ggml model (default /app/models/ggml-base.bin, matches the image's baked-in WHISPER_MODEL build arg). Point at a different mounted model file to switch without rebuilding.
  • WHISPER_LANGUAGE — default ISO-639-1 language when a request omits one (default auto = per-request detection). Pinning a language skips detection and can be faster/more accurate for single-language deployments.
  • WHISPER_THREADS — pin CPU thread count per transcription job (default: auto-detected). Useful to cap CPU usage on a small/shared VPS, e.g. WHISPER_THREADS=1.
  • WHISPER_MAX_CONCURRENT — concurrent transcription jobs allowed (default 1). Requests beyond this get 429.
⁠Request concurrency and vision limits

How many chat requests may run at once comes from your plan (max_concurrent_requests) and is separate from how many devices the licence may be installed on. Requests over the limit queue for a free slot rather than failing, so a burst of webhook deliveries drains at the licensed rate.

  • MAX_CONCURRENT_REQUESTS — local ceiling; can only lower the plan limit, never raise it.
  • QUEUE_WAIT_MS — how long an over-limit request waits before 429 (default 60000; 0 restores immediate rejection). Keep it below your caller's own HTTP timeout.
  • MAX_QUEUE_DEPTH — parked requests before shedding load (default 100).
  • MAX_IMAGE_FILES — images per request (default 10).
  • MAX_IMAGE_BYTES — per-image ceiling (default 10485760, i.e. 10MB).
  • MAX_BODY_BYTES — JSON body limit for /v1/chat/*. Derived from the image limits when unset. A 4MB camera photo becomes ~5.5MB of base64, so this must exceed it.
  • MAX_DEFAULT_BODY_BYTES — body limit for all other routes (default 1048576).

If a reverse proxy fronts the bridge, raise its limit too — nginx's client_max_body_size defaults to 1MB and returns its own 413 first.

429 responses from the concurrency limit carry a Retry-After header; honour it rather than retrying immediately.

  • WHISPER_TRANSCRIBE_TIMEOUT_MS — timeout for the actual transcription step (default 600000 / 10 min), separate from WHISPER_TIMEOUT_MS (decode only). Raise this if larger models on slower hardware still time out.
  • MAX_AUDIO_BYTES — max upload size in bytes (default 26214400 / 25 MB, matching OpenAI's limit)
  • WHISPER_BIN, FFMPEG_BIN, WHISPER_TIMEOUT_MS, WHISPER_HEALTH_TIMEOUT_MS — advanced/rarely need changing; see docs/config.md

Also supported — input/output size limits (MAX_*, SUMMARY_THRESHOLD_TOKENS), token counting (TIKTOKEN_ENCODING, IMAGE_TOKEN_*), the insights scheduler (INSIGHTS_*), SMTP delivery for insight reports and CLI update alerts (SMTP_*, INSIGHTS_EMAIL_ENABLED, CLI_UPDATE_EMAIL_ENABLED), CLI version update checks (CLI_UPDATE_*), and watchdog self-healing (WATCHDOG_*). These are optional/advanced and safe to omit for a first deployment.

Reverse-proxy vars like VIRTUAL_HOST, VIRTUAL_PORT, LETSENCRYPT_HOST, LETSENCRYPT_EMAIL are not read by CLIBridge — they're for an nginx-proxy + letsencrypt-companion sidecar, if you run one alongside this container.

⁠API call example

curl http://localhost:3900/v1/chat/completions \
  -H "Authorization: Bearer $BRIDGE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "messages": [{ "role": "user", "content": "Hello!" }]
  }'

⁠HTTP endpoints

All routes below require BRIDGE_TOKEN (Authorization: Bearer <token> or X-Bridge-Token: <token>) except GET /. Everything except /v1/license/* also requires an activated license.

OpenAI-compatible

MethodPathDescription
GET/v1/modelsList available models
POST/v1/models/refreshProbe the CLI and refresh the model cache
POST/v1/chat/completionsChat completions — JSON or SSE streaming
POST/v1/chat/completions/uploadChat completions with image upload(s) (multipart)
POST/v1/audio/transcriptionsAudio transcription — 100% local whisper.cpp, never sent externally (multipart)
POST/v1/audio/cancelCancel in-progress transcription(s), freeing capacity immediately
GET/v1/audio/statusWhisper backend status — binary/model/ffmpeg health, capacity, usage metrics
GET/v1/audio/modelsWhisper model catalog — downloaded/active state, live download progress
POST/v1/audio/models/:name/selectSwitch model, downloading it first if needed
POST/v1/audio/models/cancelCancel an in-progress model download
DELETE/v1/audio/models/:nameDelete a downloaded model to free disk space

CLI auth

MethodPathDescription
GET/v1/auth/codex, /v1/auth/claudeAuth status for each backend
POST/v1/auth/codex/start, /v1/auth/claude/startStart browser login
POST/v1/auth/claude/:id/codeSubmit Claude's verification code
GET/v1/auth/sessions/:idPoll a login session's status
POST/v1/auth/codex/update, /v1/auth/claude/updateUpdate the CLI in the running container
DELETE/v1/auth/codex, /v1/auth/claudeLog out / clear stored credentials

License

MethodPathDescription
GET/v1/license/statusCurrent license state
POST/v1/license/session/startStart device activation
POST/v1/license/refreshForce a live license check
POST/v1/license/revokeRevoke the current session

Diagnostics

MethodPathDescription
GET/v1/healthBasic health; ?details=true for CLI/watchdog/license status
GET/v1/metricsLive request/token counters
GET/v1/watchdog/statusWatchdog config and unhealthy-check count
GET/v1/logsRecent operational log entries (chat/streaming request failures, transcription/download failures, timeouts) — resets on restart
GET/v1/insights/latest, /v1/insights/historyGenerated insight reports
POST/v1/insights/generateTrigger an out-of-schedule insight report

Full request/response shapes: see docs/endpoints.md in the repo, or Swagger UI at /api-docs on your running instance (/v1/license/* and GET / are intentionally excluded from Swagger).

⁠Important notes

  • Auth state is stored in the Docker volume mounted at /data
  • Do not use docker compose down -v unless you intentionally want to erase activation and CLI auth state

Use 2.3.0 or newer.

Tag summary

Content type

Image

Digest

sha256:69cc94490…

Size

459.5 MB

Last updated

about 1 month ago

docker pull thebuildguild/cli-bridge