Clients connect to one AI Gateway; worker servers handle inference behind it.
207
The AI Gateway is the farm front-end: apps connect to it exactly as if it were a single AI Server
(same /v1/* endpoints, same Bearer key, same policy/TOFU fingerprint), and it transparently proxies every
data-plane request — chat, embeddings, audio (including the live-STT WebSocket), images, vision, Ollama-compat
/api/* — to a pool of worker AI Servers, drain-aware and with failover. No client app changes, ever —
that property is the product. Design + gap analysis: docs/v2/architecture/ai-gateway-design.md.
It is the same aisuite-server binary as deploy/docker/aiserver/, run with AISUITE_ENGINE=gateway —
one image family, two roles. All the server container's AISUITE_* env applies (see that README); extras:
| Var | Default | Meaning |
|---|---|---|
AISUITE_ENGINE | gateway | Selects the Gateway role (transparent proxy data plane). |
AISUITE_GATEWAY_ACTIVITY_SECONDS | 600 | Idle-activity timeout per forwarded request (streaming resets it). |
AISUITE_GATEWAY_RETRY_AFTER_SECONDS | 5 | Retry-After advertised on 503 (overload / all workers down). |
/data/gateway.json{
"enabled": true,
"routing_strategy": "least-latency", // round-robin (default) | least-latency | least-connections | weighted | sticky
"allow_untrusted_worker_tls": false, // accept private-CA/self-signed worker HTTPS (internal farms only)
"failover_retries": 3,
"circuit_breaker_threshold": 3,
"circuit_breaker_cooldown_seconds": 30,
"health_check_seconds": 15,
"workers": [
{ "id": "gpu-1", "base_url": "http://10.0.0.11:8080", "bearer_token": "<key issued ON that worker>", "weight": 2, "max_in_flight": 8 },
{ "id": "gpu-2", "base_url": "http://10.0.0.12:8080", "bearer_token": "<key issued ON that worker>" }
]
}
round-robin (even spread), least-latency (EWMA response time),
least-connections (fewest live requests — best when request costs vary wildly), weighted
(bias by per-worker weight), or sticky (rendezvous-hash a caller's install id / key / IP to a
preferred worker so repeat callers keep hitting the same warm model / prompt cache; failover cascades
deterministically). Per-worker max_in_flight caps concurrency: saturated workers are skipped and a
fully saturated pool answers 503 + Retry-After — backpressure, not queue collapse./v1/models/installed on the health cadence) and steers a request to a worker that already has the
requested model warm — no accidental multi-GB cold pull on an arbitrary box. GET /v1/models /
/v1/models/installed / /api/models / /api/tags on the Gateway answer with the FARM's aggregate
catalog (union across workers, served from cache)./readyz — a draining worker (rolling upgrade) is pulled
from routing before it dies. Pre-readiness workers (no /readyz) fall back to /api/health automatically.X-Request-Id) reused across attempts; deterministic 4xx/500 pass through verbatim.
All-down ⇒ 503 + Retry-After, which shipped clients already honor.bearer_token. Identity headers (app/install/machine) pass through so worker audit attributes the real
caller; run workers with AISUITE_TRUST_PROXY=1 so X-Forwarded-For is honored for their per-IP limits."auto_discover_workers": true +
"discovered_worker_token"), useful on-prem; in cloud networks list workers statically.docker build -f deploy/docker/aigateway/Dockerfile -t aisuite-gateway:local .
docker run --rm -p 8080:8080 -v gateway-data:/data aisuite-gateway:local
Like any non-loopback AI Server bind, the Gateway fail-closes without a Pro entitlement + at least one
Gateway API key under /data (licensing/server-entitlement.json, credentials/server-keys.json) — those
are the keys clients present to the Gateway; worker keys live in gateway.json.
# Gateway deployment: probe /livez (restart) + /readyz (routing), like any P1 server pod.
livenessProbe: { httpGet: { path: /livez, port: 8080 } }
readinessProbe: { httpGet: { path: /readyz, port: 8080 } }
terminationGracePeriodSeconds: 40
# Workers: a headless Service or static addresses in gateway.json; rolling upgrades drain via /readyz,
# and the Gateway routes around draining pods automatically — zero-downtime end to end.
GET /v1/gateway/workers (authenticated) — live per-worker health, breaker, EWMA latency, in-flight,
totals, catalog size, canary flag.X-AISuite-Backend: <worker id>; the edge
audit records worker_id; GET /v1/server/usage?by=worker slices fleet usage per backend — caller ×
worker × model × latency from the Gateway's single vantage."canary": true and set pool "canary_percent": 10 — a
caller-consistent 10% slice prefers the canary (new model/daemon build); everyone else stays on
stable; either group is the other's failover.GET /v1/server/metrics — the Gateway's own Prometheus metrics; GET /v1/server/usage — edge usage
rollups (per app/install/key — FinOps attribution rides the forwarded identity headers).Content type
Image
Digest
sha256:e3108d5cf…
Size
706.1 MB
Last updated
1 day ago
docker pull softwaretailor/aigateway