Sign inSign up

softwaretailor/aigateway

By softwaretailor

•Updated 1 day ago

Clients connect to one AI Gateway; worker servers handle inference behind it.

Image
0

207

softwaretailor/aigateway repository overview

⁠AI Gateway in a container

The AI Gateway is the farm front-end: apps connect to it exactly as if it were a single AI Server (same /v1/* endpoints, same Bearer key, same policy/TOFU fingerprint), and it transparently proxies every data-plane request — chat, embeddings, audio (including the live-STT WebSocket), images, vision, Ollama-compat /api/* — to a pool of worker AI Servers, drain-aware and with failover. No client app changes, ever — that property is the product. Design + gap analysis: docs/v2/architecture/ai-gateway-design.md.

It is the same aisuite-server binary as deploy/docker/aiserver/, run with AISUITE_ENGINE=gateway — one image family, two roles. All the server container's AISUITE_* env applies (see that README); extras:

VarDefaultMeaning
AISUITE_ENGINEgatewaySelects the Gateway role (transparent proxy data plane).
AISUITE_GATEWAY_ACTIVITY_SECONDS600Idle-activity timeout per forwarded request (streaming resets it).
AISUITE_GATEWAY_RETRY_AFTER_SECONDS5Retry-After advertised on 503 (overload / all workers down).

⁠Worker pool — /data/gateway.json

{
  "enabled": true,
  "routing_strategy": "least-latency",      // round-robin (default) | least-latency | least-connections | weighted | sticky
  "allow_untrusted_worker_tls": false,       // accept private-CA/self-signed worker HTTPS (internal farms only)
  "failover_retries": 3,
  "circuit_breaker_threshold": 3,
  "circuit_breaker_cooldown_seconds": 30,
  "health_check_seconds": 15,
  "workers": [
    { "id": "gpu-1", "base_url": "http://10.0.0.11:8080", "bearer_token": "<key issued ON that worker>", "weight": 2, "max_in_flight": 8 },
    { "id": "gpu-2", "base_url": "http://10.0.0.12:8080", "bearer_token": "<key issued ON that worker>" }
  ]
}
  • Load balancing: round-robin (even spread), least-latency (EWMA response time), least-connections (fewest live requests — best when request costs vary wildly), weighted (bias by per-worker weight), or sticky (rendezvous-hash a caller's install id / key / IP to a preferred worker so repeat callers keep hitting the same warm model / prompt cache; failover cascades deterministically). Per-worker max_in_flight caps concurrency: saturated workers are skipped and a fully saturated pool answers 503 + Retry-After — backpressure, not queue collapse.
  • Model-aware routing: the Gateway learns each worker's installed models (polled from /v1/models/installed on the health cadence) and steers a request to a worker that already has the requested model warm — no accidental multi-GB cold pull on an arbitrary box. GET /v1/models / /v1/models/installed / /api/models / /api/tags on the Gateway answer with the FARM's aggregate catalog (union across workers, served from cache).
  • Health/drain: each worker is probed on /readyz — a draining worker (rolling upgrade) is pulled from routing before it dies. Pre-readiness workers (no /readyz) fall back to /api/health automatically.
  • Failover: only on the transient signals (connect failure, upstream 502/503/504) with the request's idempotency id (X-Request-Id) reused across attempts; deterministic 4xx/500 pass through verbatim. All-down ⇒ 503 + Retry-After, which shipped clients already honor.
  • Trust boundary: a client's Gateway API key never reaches a worker — the Gateway presents each worker's bearer_token. Identity headers (app/install/machine) pass through so worker audit attributes the real caller; run workers with AISUITE_TRUST_PROXY=1 so X-Forwarded-For is honored for their per-IP limits.
  • LAN worker auto-discovery (mDNS) is supported ("auto_discover_workers": true + "discovered_worker_token"), useful on-prem; in cloud networks list workers statically.

⁠Run

docker build -f deploy/docker/aigateway/Dockerfile -t aisuite-gateway:local .
docker run --rm -p 8080:8080 -v gateway-data:/data aisuite-gateway:local

Like any non-loopback AI Server bind, the Gateway fail-closes without a Pro entitlement + at least one Gateway API key under /data (licensing/server-entitlement.json, credentials/server-keys.json) — those are the keys clients present to the Gateway; worker keys live in gateway.json.

⁠Kubernetes sketch (gateway + workers)
# Gateway deployment: probe /livez (restart) + /readyz (routing), like any P1 server pod.
livenessProbe:  { httpGet: { path: /livez,  port: 8080 } }
readinessProbe: { httpGet: { path: /readyz, port: 8080 } }
terminationGracePeriodSeconds: 40
# Workers: a headless Service or static addresses in gateway.json; rolling upgrades drain via /readyz,
# and the Gateway routes around draining pods automatically — zero-downtime end to end.

⁠Observability

  • GET /v1/gateway/workers (authenticated) — live per-worker health, breaker, EWMA latency, in-flight, totals, catalog size, canary flag.
  • Per-worker attribution: every served response carries X-AISuite-Backend: <worker id>; the edge audit records worker_id; GET /v1/server/usage?by=worker slices fleet usage per backend — caller × worker × model × latency from the Gateway's single vantage.
  • Canary rollout: flag a worker "canary": true and set pool "canary_percent": 10 — a caller-consistent 10% slice prefers the canary (new model/daemon build); everyone else stays on stable; either group is the other's failover.
  • GET /v1/server/metrics — the Gateway's own Prometheus metrics; GET /v1/server/usage — edge usage rollups (per app/install/key — FinOps attribution rides the forwarded identity headers).
  • The Gateway writes the standard content-free audit JSONL at the edge — a single vantage point over the whole farm (per-worker token counts remain on each worker's own audit).

Tag summary

Content type

Image

Digest

sha256:e3108d5cf…

Size

706.1 MB

Last updated

1 day ago

docker pull softwaretailor/aigateway