Lightweight LLM API gateway with multi-provider routing, rate limiting, and failover
1.8K
Lightweight LLM API Gateway — route, rate-limit, cache, and track usage across OpenAI, Ollama, Anthropic, and more.
Running multiple LLM providers? ai-gateway sits in front of all of them and gives you:
docker run -d \
--name ai-gateway \
-p 8080:8080 \
-e OPENAI_API_KEY=sk-... \
-v ./config.yaml:/app/config.yaml \
zachbg/ai-gateway
Then use it like OpenAI:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello!"}]
}'
services:
ai-gateway:
image: zachbg/ai-gateway
ports:
- "8080:8080"
volumes:
- ./config.yaml:/app/config.yaml
- gateway-data:/app/data
environment:
- OPENAI_API_KEY=${OPENAI_API_KEY}
- ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY}
ollama:
image: ollama/ollama
volumes:
- ollama-data:/root/.ollama
volumes:
gateway-data:
ollama-data:
Copy config.example.yaml and customize:
providers:
openai:
base_url: "https://api.openai.com/v1"
api_key: "${OPENAI_API_KEY}"
models: ["gpt-4o", "gpt-4o-mini"]
ollama:
base_url: "http://ollama:11434"
models: ["llama3.3", "mistral"]
rate_limit:
enabled: true
requests_per_minute: 60
cache:
enabled: true
ttl_seconds: 3600
| Method | Path | Description |
|---|---|---|
POST | /v1/chat/completions | Proxy chat completion (OpenAI-compatible) |
GET | /v1/models | List all available models |
GET | /v1/usage | View recent usage stats |
GET | /health | Health check |
docker buildx build --platform linux/amd64,linux/arm64 -t zachbg/ai-gateway --push .
MIT
Content type
Image
Digest
sha256:ab92bebfc…
Size
45.3 MB
Last updated
6 months ago
docker pull zachbg/ai-gateway