HTTP AI gateway
1.3K
Pass-through AI gateway: one Bearer key fans out to OpenAI / Responses / Gemini / Claude / embeddings upstreams. No request translation, no keys baked into the image.
Two ways: A) copy this folder (files included) — fastest. B) start from scratch — every file is printed in full below, paste and run.
cp .env.example .env # set POSTGRES_PASSWORD
cp keys.example.json keys.json # real apiKeys
cp budgets.example.json budgets.json
mkdir -p data
docker compose up -d
curl http://localhost:8787/health # {"ok":true,...}
Create the files below side by side, then mkdir -p data && docker compose up -d.
docker-compose.ymlname: my-aigateway
services:
gateway:
image: kienxuandaoit/my_aigateway:${GW_IMAGE_TAG:-latest}
ports:
# GW_HOST=127.0.0.1 for loopback-only (TLS proxy on same host).
# Container side stays 8787 (app PORT default; change both if overridden).
- "${GW_HOST:-0.0.0.0}:${GW_PORT:-8787}:8787"
env_file:
- path: .env
required: false
environment:
# compose network hostname. overrides GW_DB from .env (usually a host sqlite path).
GW_DB: postgres://postgres:${POSTGRES_PASSWORD:-secret}@db:5432/gw
volumes:
- ./keys.json:/app/keys.json:ro
# presets + pricing catalog are standalone files (not baked into the
# image): edit here, restart to apply. regen catalog: ./gen-models.sh
- ./presets.json:/app/presets.json:ro
- ./models.json:/app/models.json:ro
# must exist on the host before `up`: Docker binds a missing path as a
# *directory*, which the app refuses to read. `cp budgets.example.json budgets.json`.
- ./budgets.json:/app/budgets.json:ro
- ./data:/app/data
restart: unless-stopped
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:8787/health"]
interval: 10s
timeout: 3s
retries: 3
start_period: 10s
depends_on:
db:
condition: service_healthy
db:
image: postgres:16-alpine
environment:
POSTGRES_DB: gw
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:-secret}
volumes:
- pgdata:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]
interval: 2s
timeout: 3s
retries: 20
# not published — host already has :5432. query via `docker compose exec db psql`
volumes:
pgdata:
.envImage tag + db password. Gateway GW_* vars go here too.
# cp .env.example .env
# Image tag to run. Must match a tag pushed to Docker Hub
# (./scripts/docker-push.sh [tag] pushes package.json version by default).
GW_IMAGE_TAG=0.1.0
# Postgres password for the bundled db service. Change it.
POSTGRES_PASSWORD=secret
# Gateway GW_* vars go here too (compose env_file passes them through).
# Full list with defaults: ../.env.example
# GW_UI_BASIC=ui:REPLACE-ME
# GW_ROLLUP=1h,24h
# GW_RETENTION=30d
keys.jsonVirtual keys (left side = Bearer token your clients send) → upstream routes. One entry per sdk. Replace every replace-me with a real key.
{
"gw-openai-completions-replace-me": {
"team": "core",
"member": "pool",
"routes": [
{
"sdk": "openai-completions",
"provider": "google",
"apiKey": "AIza-replace-me",
"model": "gemini-3.5-flash-lite"
}
]
},
"gw-openai-responses-replace-me": {
"team": "core",
"member": "pool",
"routes": [
{
"sdk": "openai-responses",
"provider": "bedrock-runtime.us-east-1",
"apiKey": "replace-me",
"model": "global.openai.gpt-5.6-sol"
}
]
},
"gw-google-replace-me": {
"team": "core",
"member": "pool",
"routes": [
{
"sdk": "google-generative-ai",
"provider": "google",
"apiKey": "AIza-replace-me",
"model": "gemini-3.5-flash-lite"
}
]
},
"gw-anthropic-messages-replace-me": {
"team": "core",
"member": "pool",
"routes": [
{
"sdk": "anthropic-messages",
"provider": "bedrock-runtime.us-east-1",
"apiKey": "replace-me",
"model": "anthropic.claude-sonnet-5"
}
]
},
"gw-embeddings-replace-me": {
"team": "core",
"member": "pool",
"routes": [
{
"sdk": "openai-embeddings",
"provider": "ollama",
"apiKey": "replace-me",
"model": "nomic-embed-text"
}
]
}
}
presets.jsonProvider routing (hosts + per-wire paths). Rarely edited; add a block to point a provider name at another host. Missing file = loud boot error.
{
"google": {
"baseURL": "https://generativelanguage.googleapis.com",
"wires": {
"openai-completions": { "path": "v1beta/openai/chat/completions", "exclude": ["store"] },
"google-generative-ai": { "path": "v1beta" }
}
},
"openrouter": {
"baseURL": "https://openrouter.ai/api",
"wires": { "openai-completions": {} }
},
"logfare": {
"baseURL": "https://logfare.ai",
"wires": { "openai-completions": {} }
},
"poolside": {
"baseURL": "https://inference.poolside.ai",
"wires": { "openai-completions": {} }
},
"ollama": {
"baseURL": "http://127.0.0.1:11434",
"wires": {
"openai-completions": {},
"openai-embeddings": { "path": "v1/embeddings" }
}
},
"opencode": {
"baseURL": "https://opencode.ai/inference",
"wires": {
"openai-completions": { "path": "openai/v1/chat/completions" },
"google-generative-ai": { "path": "google/v1beta" }
}
},
"openai": {
"baseURL": "https://api.openai.com",
"wires": {
"openai-completions": {},
"openai-responses": {},
"openai-embeddings": {}
}
},
"anthropic": {
"baseURL": "https://api.anthropic.com",
"wires": { "anthropic-messages": {} }
},
"deepseek": {
"baseURL": "https://api.deepseek.com",
"wires": { "openai-completions": {} }
},
"xai": {
"baseURL": "https://api.x.ai",
"wires": { "openai-completions": {} }
},
"mistral": {
"baseURL": "https://api.mistral.ai",
"wires": { "openai-completions": {} }
},
"groq": {
"baseURL": "https://api.groq.com/openai",
"wires": { "openai-completions": {} }
},
"bedrock-runtime": {
"baseURL": "https://bedrock-runtime.{region}.amazonaws.com",
"wires": {
"openai-completions": { "path": "openai/v1/chat/completions" },
"openai-responses": { "path": "openai/v1/responses" },
"anthropic-messages": { "path": "anthropic/v1/messages" }
}
},
"bedrock-mantle": {
"baseURL": "https://bedrock-mantle.{region}.api.aws",
"wires": {
"openai-completions": {},
"openai-responses": { "path": "v1/responses" },
"anthropic-messages": { "path": "anthropic/v1/messages" }
}
}
}
models.jsonPrice catalog per serving path (provider/model, exactly as on the wire). Missing file = every est_cost null, spend caps silently never fire. Don't hand-write new pins — add them to keys.json, then regen:
cd docker # gen-models.sh lives beside keys.json
./gen-models.sh # tag from .env GW_IMAGE_TAG
./gen-models.sh 0.1.0 # ...or pick a tag explicitly
What it needs: keys.json with real pins (only provider/model are read, apiKeys unused), an existing models.json file (from-scratch: echo '{}' > models.json first — Docker binds a missing path as a directory), and network to models.dev. Host needs no node: fill runs as dist/models.js --fill inside the gateway image. Output per pin: fill (new entry), fill meta, refresh, manual-diff (hand-verified price kept, verify drift yourself), unmapped (no catalog match — add by hand, exit 1). Then docker compose restart gateway to apply.
Starter content:
{
"google/gemini-3.5-flash-lite": {
"input": 0.3,
"output": 2.5,
"cache_read": 0.03,
"cache_write": null,
"tiers": [],
"meta": {
"context_length": 1048576,
"max_completion_tokens": 65536,
"input_modalities": [
"text",
"image",
"video",
"audio",
"pdf"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "google/gemini-3.5-flash-lite",
"updatedAt": "2026-09-28T07:00:43.195Z"
},
"opencode/gemini-3.5-flash-lite": {
"input": 0.3,
"output": 2.5,
"cache_read": 0.03,
"cache_write": null,
"tiers": [],
"meta": {
"context_length": 1048576,
"max_completion_tokens": 65536,
"input_modalities": [
"text",
"image",
"video",
"audio",
"pdf"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "opencode/gemini-3.5-flash-lite",
"updatedAt": "2026-09-28T07:00:43.195Z"
},
"opencode/muse-spark-1.3": {
"input": 1.25,
"output": 4.25,
"cache_read": 0.15,
"cache_write": null,
"tiers": [],
"meta": {
"context_length": 1048576,
"max_completion_tokens": 131072,
"input_modalities": [
"text",
"image",
"video",
"pdf",
"audio"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "opencode/muse-spark-1.3",
"updatedAt": "2026-09-28T12:44:51.912Z"
},
"deepseek/deepseek-flash": {
"input": 0.15,
"output": 0.6,
"cache_read": 0.003,
"cache_write": null,
"tiers": [],
"meta": {
"context_length": 1000000,
"max_completion_tokens": 393216,
"input_modalities": [
"text",
"image"
],
"output_modalities": [
"text"
]
},
"source": "manual",
"ref": "https://api-docs.deepseek.com/quick_start/pricing (off-peak base, peak 01-04+06-10 UTC Mon-Fri x2)",
"schedule": {
"peakMult": 2,
"days": [
1,
2,
3,
4,
5
],
"hours": [
[
1,
4
],
[
6,
10
]
]
},
"updatedAt": "2026-09-28T14:00:00.000Z"
},
"openrouter/openrouter/free": {
"input": 0,
"output": 0,
"cache_read": null,
"cache_write": null,
"tiers": [],
"meta": {
"context_length": 200000,
"max_completion_tokens": 8000,
"input_modalities": [
"text",
"image"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "openrouter/openrouter/free",
"updatedAt": "2026-09-30T04:06:49.384Z"
},
"opencode/big-pickle": {
"input": 0,
"output": 0,
"cache_read": 0,
"cache_write": 0,
"tiers": [],
"meta": {
"context_length": 200000,
"max_completion_tokens": 32000,
"input_modalities": [
"text"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "opencode/big-pickle",
"updatedAt": "2026-09-30T04:06:49.386Z"
},
"bedrock-runtime.us-east-1/global.openai.gpt-5.6-sol": {
"input": 4,
"output": 20,
"cache_read": 0.4,
"cache_write": 5,
"tiers": [
{
"size": 272000,
"input": 8,
"output": 30,
"cache_read": 0.8
}
],
"meta": {
"context_length": 1050000,
"max_completion_tokens": 128000,
"input_modalities": [
"text",
"image"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "amazon-bedrock/global.openai.gpt-5.6-sol",
"updatedAt": "2026-09-30T04:06:49.387Z"
},
"bedrock-mantle.us-east-1/openai.gpt-oss-120b": {
"input": 0.15,
"output": 0.6,
"cache_read": null,
"cache_write": null,
"tiers": [],
"meta": {
"context_length": 131072,
"max_completion_tokens": 131072,
"input_modalities": [
"text"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "amazon-bedrock/openai.gpt-oss-120b",
"updatedAt": "2026-09-30T04:06:49.388Z"
},
"bedrock-runtime.us-east-1/anthropic.claude-sonnet-5": {
"input": 2,
"output": 10,
"cache_read": 0.2,
"cache_write": 2.5,
"tiers": [],
"meta": {
"context_length": 1000000,
"max_completion_tokens": 128000,
"input_modalities": [
"text",
"image",
"pdf"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "amazon-bedrock/anthropic.claude-sonnet-5",
"updatedAt": "2026-09-30T04:06:49.388Z"
},
"bedrock-mantle.eu-central-1/anthropic.claude-sonnet-5": {
"input": 2,
"output": 10,
"cache_read": 0.2,
"cache_write": 2.5,
"tiers": [],
"meta": {
"context_length": 1000000,
"max_completion_tokens": 128000,
"input_modalities": [
"text",
"image",
"pdf"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "amazon-bedrock/anthropic.claude-sonnet-5",
"updatedAt": "2026-09-30T04:06:49.389Z"
},
"ollama/nomic-embed-text": {
"input": 0.05,
"output": 0,
"cache_read": null,
"cache_write": null,
"tiers": [],
"meta": {
"context_length": 8192,
"max_completion_tokens": 768,
"input_modalities": [
"text"
],
"output_modalities": [
"text"
]
},
"source": "models.dev",
"ref": "nomic-embed-text",
"updatedAt": "2026-09-30T04:06:49.390Z"
}
}
budgets.jsonSymbolic spend caps (reference estimates, not bills). Empty team/member/model = wildcard. Must exist before up.
{
"_comment": "cp budgets.example.json budgets.json. Symbolic caps on reference estimates (est_cost), not bills. Empty team/member/model = wildcard. alerts fire via rollup tick (log) + /ui pill. \"enforce\": true additionally rejects requests past the cap with 429; \"concurrency\" (needs enforce) caps in-flight requests for the scope.",
"budgets": [
{ "team": "core", "member": "pool", "window": "24h", "usd": 10, "enforce": true, "concurrency": 4 },
{ "team": "core", "member": "pool", "model": "flash", "window": "7d", "usd": 20 }
]
}
curl http://localhost:8787/health → {"ok":true,...}.curl -X POST http://localhost:8787/v1/chat/completions -H "Authorization: Bearer <your-token>" -H "Content-Type: application/json" -d '{"model":"<pin>","messages":[{"role":"user","content":"ping"}]}'.docker compose exec db psql -U postgres -d gw.GW_IMAGE_TAG in .env, docker compose up -d.GW_HOST=127.0.0.1..env baked into the image. Secrets stay mounts/env.Content type
Image
Digest
sha256:c52f69ca1…
Size
59 MB
Last updated
2 days ago
docker pull kienxuandaoit/my_aigateway