Chief Operating Officer for Our Large Language Model Providers
1.4K
๐ Intelligent Load Balancer for LLM APIs with Full OpenAI Compatibility
COO-LLM is a high-performance reverse proxy that intelligently distributes requests across multiple LLM providers (OpenAI, Google Gemini, Anthropic Claude) and API keys. It provides seamless OpenAI API compatibility, advanced load balancing algorithms, real-time cost optimization, and enterprise-grade observability.
docker run -p 2906:2906 \
-e OPENAI_API_KEY="sk-your-key" \
-v $(pwd)/configs:/app/configs \
khapu2906/coo-llm:latest
Docker Hub Images:
khapu2906/coo-llm:latest - Latest development buildkhapu2906/coo-llm:v1.0.2 - Specific version tagsCOO-LLM works seamlessly with LangChain and other OpenAI-compatible libraries:
// JavaScript/TypeScript
import { ChatOpenAI } from '@langchain/openai';
const llm = new ChatOpenAI({
modelName: 'gpt-4o',
openAIApiKey: 'dummy-key', // Ignored by COO-LLM
configuration: {
baseURL: 'http://localhost:2906/v1',
},
});
// Simple request
const response = await llm.invoke('Hello!');
// Conversation history
const messages = [
new HumanMessage('What is AI?'),
new AIMessage('AI stands for Artificial Intelligence...'),
new HumanMessage('How does it work?'),
];
const response = await llm.invoke(messages);
# Python
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="gpt-4o",
openai_api_key="dummy-key", # Ignored by COO-LLM
openai_api_base="http://localhost:2906/v1"
)
response = llm.invoke("Hello!")
print(response.content)
See langchain-demo/โ for complete examples.
Complete documentation is available in the docsโ directory.
COO-LLM uses YAML configuration with environment variable support:
version: "1.0"
# Server configuration
server:
listen: ":2906"
admin_api_key: "${ADMIN_KEY}"
# Logging configuration
logging:
file:
enabled: true
path: "./logs/coo-llm.log"
max_size_mb: 100
prometheus:
enabled: true
endpoint: "/metrics"
# LLM Providers configuration
llm_providers:
- id: "openai-prod"
type: "openai"
api_keys: ["${OPENAI_KEY_1}", "${OPENAI_KEY_2}"]
base_url: "https://api.openai.com"
model: "gpt-4o"
pricing:
input_token_cost: 0.002
output_token_cost: 0.01
limits:
req_per_min: 200
tokens_per_min: 100000
- id: "gemini-prod"
type: "gemini"
api_keys: ["${GEMINI_KEY_1}"]
base_url: "https://generativelanguage.googleapis.com"
model: "gemini-1.5-pro"
pricing:
input_token_cost: 0.00025
output_token_cost: 0.0005
limits:
req_per_min: 150
tokens_per_min: 80000
# API Key permissions (optional - if not specified, all keys have full access)
api_keys:
- key: "client-a-key"
allowed_providers: ["openai-prod"] # Only OpenAI access
description: "Client A - OpenAI only"
- key: "premium-key"
allowed_providers: ["openai-prod", "gemini-prod"] # Full access
description: "Premium client with all providers"
- key: "test-key"
allowed_providers: ["*"] # Wildcard for all providers
description: "Development key"
# Model aliases for easy reference (maps to provider_id:model)
model_aliases:
gpt-4o: openai-prod:gpt-4o
gemini-pro: gemini-prod:gemini-1.5-pro
claude-opus: claude-prod:claude-3-opus
# Load balancing policy
policy:
algorithm: "hybrid" # "round_robin", "least_loaded", "hybrid"
priority: "balanced" # "balanced", "cost", "req", "token"
retry:
max_attempts: 3
timeout: "30s"
interval: "1s"
cache:
enabled: true
ttl_seconds: 10
# Storage configuration
storage:
runtime:
type: "redis" # "memory", "redis", "file", "http"
addr: "localhost:6379"
password: "${REDIS_PASSWORD}"
See Configuration Guideโ for complete options.
COO-LLM implements enterprise-grade security measures to protect your LLM API infrastructure:
Client Authentication: Configure API keys with granular permissions:
# In config.yaml
api_keys:
- key: "client-a-key"
allowed_providers: ["openai-prod"] # Only OpenAI access
description: "Client A limited access"
- key: "premium-key"
allowed_providers: ["openai-prod", "gemini-prod"] # Full access
description: "Premium client"
- key: "test-key"
allowed_providers: ["*"] # Wildcard for all providers
description: "Development key"
Usage: Include the API key in the Authorization header:
curl -X POST http://localhost:2906/v1/chat/completions \
-H "Authorization: Bearer your-secure-api-key-1" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'
The admin API (/admin/*) requires additional authentication:
server:
admin_api_key: "your-admin-secret"
Access admin endpoints:
curl -H "Authorization: Bearer your-admin-secret" \
http://localhost:2906/admin/v1/config
For production deployments:
COO-LLM provides 100% OpenAI API compatibility:
POST /v1/chat/completions - Chat completions with conversation historyGET /v1/models - List available modelsPOST /admin/v1/config/validate - Config validation (admin)GET /admin/v1/config - Get current config (admin)GET /metrics - Prometheus metricsWe welcome contributions! Please see our Contributing Guidelinesโ for details.
git clone https://github.com/your-org/coo-llm.git
cd coo-llm
go mod download
go build ./...
go test ./...
This project is licensed under the DIB License v1.0 - see the LICENSEโ file for details.
COO-LLM - The Intelligent LLM API Load Balancer ๐
Load balance your LLM API calls across multiple providers with OpenAI compatibility, real-time cost optimization, and enterprise-grade reliability.
Content type
Image
Digest
sha256:99e808b2fโฆ
Size
46.9 MB
Last updated
12 months ago
docker pull khapu2906/coo-llm