GoModel - high-performance, lightweight AI gateway written in Go
GoModel is a fast, self-hosted AI gateway and LLM proxy with OpenAI-compatible and Anthropic-compatible APIs. Connect your applications to one endpoint and route requests across cloud and local AI model providers.
It is a lightweight LiteLLM alternative for teams that want lower AI costs, reliable model access, and complete LLM observability without a heavy control plane.
GoModel saves you money and nerves.
Money - cache repeated requests, track every token and cost, enforce budgets, and route traffic to the right model.
Nerves - keep applications running with load balancing, retries, circuit breakers, provider health checks, and automatic failover.
Multi-provider AI gateway - use OpenAI, Anthropic, Google Gemini, Vertex AI, Azure OpenAI, Amazon Bedrock, Cohere, DeepSeek, Groq, xAI, OpenRouter, Ollama, vLLM, and many more through one API.
OpenAI and Anthropic API compatibility - supports Chat Completions, Responses API, Conversations, Messages API, embeddings, audio, files, batches, and realtime APIs. Use the official SDKs by changing the base URL.
Virtual models and load balancing - expose stable model aliases and balance requests with round-robin or cost-based routing.
Failover and resilience - reroute failed requests to backup models or providers with retries, circuit breakers, and live provider health.
Response caching - reduce latency and cost with exact caching, semantic caching, and provider-native prompt caching.
Cost tracking and usage analytics - understand spend by provider, model, user path, key, or label from the built-in dashboard.
Budgets and rate limits - enforce spend limits, requests per minute, tokens per minute, and concurrency limits.
Access control - create managed API keys and scope model access, usage, budgets, workflows, and audit logs with hierarchical user paths.
Guardrails - inspect, modify, or reject requests and responses at the gateway.
MCP gateway - aggregate Model Context Protocol servers behind one authenticated endpoint with tool access controls, usage tracking, and rate limits.
Provider-native passthrough - access native provider APIs while keeping GoModel authentication, tracking, and observability.
Provider key rotation - spread traffic across multiple API keys while preserving session affinity for prompt caching.
Full observability - real-time request logs, audit logs, usage analytics, provider status, cost estimates, and Prometheus metrics.
Flexible configuration - configure GoModel with environment variables, config.yaml, or the admin dashboard without restarting.
Production-friendly container - compact distroless image, non-root runtime, built-in health check, and support for linux/amd64, linux/arm64, and linux/arm/v7.