Sign inSign up

khapu2906/coo-llm

By khapu2906

โ€ขUpdated 12 months ago

Chief Operating Officer for Our Large Language Model Providers

Image
Security
Machine learning & AI
Operating systems
0

1.4K

khapu2906/coo-llm repository overview

โ COO-LLM

๐Ÿš€ Intelligent Load Balancer for LLM APIs with Full OpenAI Compatibility

COO-LLM is a high-performance reverse proxy that intelligently distributes requests across multiple LLM providers (OpenAI, Google Gemini, Anthropic Claude) and API keys. It provides seamless OpenAI API compatibility, advanced load balancing algorithms, real-time cost optimization, and enterprise-grade observability.

Go Version Docker License: DIB OpenAI Compatible

โ ๐Ÿš€ Features

โ โœจ Core Capabilities
  • ๐Ÿ”„ Full OpenAI API Compatibility: Drop-in replacement with identical request/response formats
  • ๐ŸŒ Multi-Provider Support: OpenAI, Google Gemini, Anthropic Claude, and custom providers
  • ๐Ÿง  Intelligent Load Balancing: Advanced algorithms (Round Robin, Least Loaded, Hybrid) with real-time optimization
  • ๐Ÿ’ฌ Conversation History: Full support for multi-turn conversations and message history
โ ๐Ÿ’ฐ Cost & Performance Optimization
  • ๐Ÿ“Š Real-time Cost Tracking: Monitor and optimize API costs across all providers
  • โšก Rate Limit Management: Sliding window rate limiting with automatic key rotation
  • ๐Ÿ“ˆ Performance Monitoring: Track latency, success rates, token usage, and error patterns
  • ๐Ÿ”„ Response Caching: Configurable caching to reduce costs and improve performance
โ ๐Ÿข Enterprise-Ready
  • ๐Ÿ”Œ Extensible Architecture: Plugin system for custom providers, storage backends, and logging
  • ๐Ÿ“Š Production Observability: Prometheus metrics, structured logging, and health checks
  • โš™๏ธ Configuration Management: YAML-based configuration with environment variable support
  • ๐Ÿ”’ Security: API key masking, secure storage, and authentication controls
โ Run
docker run -p 2906:2906 \
  -e OPENAI_API_KEY="sk-your-key" \
  -v $(pwd)/configs:/app/configs \
  khapu2906/coo-llm:latest

Docker Hub Images:

  • khapu2906/coo-llm:latest - Latest development build
  • khapu2906/coo-llm:v1.0.2 - Specific version tags
โ ๐Ÿง  LangChain Integration

COO-LLM works seamlessly with LangChain and other OpenAI-compatible libraries:

// JavaScript/TypeScript
import { ChatOpenAI } from '@langchain/openai';

const llm = new ChatOpenAI({
  modelName: 'gpt-4o',
  openAIApiKey: 'dummy-key', // Ignored by COO-LLM
  configuration: {
    baseURL: 'http://localhost:2906/v1',
  },
});

// Simple request
const response = await llm.invoke('Hello!');

// Conversation history
const messages = [
  new HumanMessage('What is AI?'),
  new AIMessage('AI stands for Artificial Intelligence...'),
  new HumanMessage('How does it work?'),
];
const response = await llm.invoke(messages);
# Python
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="gpt-4o",
    openai_api_key="dummy-key",  # Ignored by COO-LLM
    openai_api_base="http://localhost:2906/v1"
)

response = llm.invoke("Hello!")
print(response.content)

See langchain-demo/โ  for complete examples.

โ ๐Ÿ“š Documentation

Complete documentation is available in the docsโ  directory.

โ Documentation Structure
  • Intro: Overview, architecture, and getting started
  • Guides: User guides, configuration, and deployment
  • Reference: Technical API, configuration, and balancer reference
  • Contributing: Development guidelines and contribution process

โ ๐Ÿ”ง Configuration

COO-LLM uses YAML configuration with environment variable support:

version: "1.0"

# Server configuration
server:
  listen: ":2906"
  admin_api_key: "${ADMIN_KEY}"

# Logging configuration
logging:
  file:
    enabled: true
    path: "./logs/coo-llm.log"
    max_size_mb: 100
  prometheus:
    enabled: true
    endpoint: "/metrics"

# LLM Providers configuration
llm_providers:
  - id: "openai-prod"
    type: "openai"
    api_keys: ["${OPENAI_KEY_1}", "${OPENAI_KEY_2}"]
    base_url: "https://api.openai.com"
    model: "gpt-4o"
    pricing:
      input_token_cost: 0.002
      output_token_cost: 0.01
    limits:
      req_per_min: 200
      tokens_per_min: 100000
  - id: "gemini-prod"
    type: "gemini"
    api_keys: ["${GEMINI_KEY_1}"]
    base_url: "https://generativelanguage.googleapis.com"
    model: "gemini-1.5-pro"
    pricing:
      input_token_cost: 0.00025
      output_token_cost: 0.0005
    limits:
      req_per_min: 150
      tokens_per_min: 80000

# API Key permissions (optional - if not specified, all keys have full access)
api_keys:
  - key: "client-a-key"
    allowed_providers: ["openai-prod"]  # Only OpenAI access
    description: "Client A - OpenAI only"
  - key: "premium-key"
    allowed_providers: ["openai-prod", "gemini-prod"]  # Full access
    description: "Premium client with all providers"
  - key: "test-key"
    allowed_providers: ["*"]  # Wildcard for all providers
    description: "Development key"

# Model aliases for easy reference (maps to provider_id:model)
model_aliases:
  gpt-4o: openai-prod:gpt-4o
  gemini-pro: gemini-prod:gemini-1.5-pro
  claude-opus: claude-prod:claude-3-opus

# Load balancing policy
policy:
  algorithm: "hybrid"  # "round_robin", "least_loaded", "hybrid"
  priority: "balanced" # "balanced", "cost", "req", "token"
  retry:
    max_attempts: 3
    timeout: "30s"
    interval: "1s"
  cache:
    enabled: true
    ttl_seconds: 10

# Storage configuration
storage:
  runtime:
    type: "redis"  # "memory", "redis", "file", "http"
    addr: "localhost:6379"
    password: "${REDIS_PASSWORD}"

See Configuration Guideโ  for complete options.

โ ๐Ÿ”’ Security

COO-LLM implements enterprise-grade security measures to protect your LLM API infrastructure:

โ API Key Authentication

Client Authentication: Configure API keys with granular permissions:

# In config.yaml
api_keys:
  - key: "client-a-key"
    allowed_providers: ["openai-prod"]  # Only OpenAI access
    description: "Client A limited access"
  - key: "premium-key"
    allowed_providers: ["openai-prod", "gemini-prod"]  # Full access
    description: "Premium client"
  - key: "test-key"
    allowed_providers: ["*"]  # Wildcard for all providers
    description: "Development key"

Usage: Include the API key in the Authorization header:

curl -X POST http://localhost:2906/v1/chat/completions \
  -H "Authorization: Bearer your-secure-api-key-1" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
โ Security Best Practices
  • ๐Ÿ” API Key Management: Rotate keys regularly and use different keys for different clients
  • ๐Ÿ“Š Access Logging: All requests are logged with client identification for audit trails
  • ๐Ÿšซ Key Masking: API keys are never logged in plain text (masked in logs and admin endpoints)
  • ๐Ÿ”’ Provider Key Security: LLM provider API keys are stored securely and never exposed
  • โšก Rate Limiting: Built-in rate limiting prevents abuse and ensures fair usage
  • ๐Ÿ›ก๏ธ Input Validation: All requests are validated before processing
โ Admin API Security

The admin API (/admin/*) requires additional authentication:

server:
  admin_api_key: "your-admin-secret"

Access admin endpoints:

curl -H "Authorization: Bearer your-admin-secret" \
  http://localhost:2906/admin/v1/config
โ Production Deployment

For production deployments:

  • Use HTTPS/TLS termination (nginx, cloud load balancer, etc.)
  • Store API keys in secure secret management systems
  • Enable audit logging and monitoring
  • Regularly update and patch the system
  • Use network security groups to restrict access

โ ๐Ÿ”— API Compatibility

COO-LLM provides 100% OpenAI API compatibility:

โ โœ… Supported Endpoints
  • POST /v1/chat/completions - Chat completions with conversation history
  • GET /v1/models - List available models
  • POST /admin/v1/config/validate - Config validation (admin)
  • GET /admin/v1/config - Get current config (admin)
  • GET /metrics - Prometheus metrics
โ โœ… Compatible Libraries
  • OpenAI SDKs: Python, Node.js, Go, etc.
  • LangChain/LangGraph: Full integration support
  • LlamaIndex: Compatible with OpenAI connector
  • Any OpenAI-compatible client
โ โœ… Features Supported
  • โœ… Conversation history (messages array)
  • โœ… Streaming responses (planned)
  • โœ… Function calling (planned)
  • โœ… Token usage tracking
  • โœ… Model aliases
  • โœ… Custom parameters (temperature, top_p, etc.)

โ ๐Ÿ“Š Key Metrics

  • ๐Ÿš€ Load Balancing: Intelligent distribution across 3+ providers
  • ๐Ÿ’ฐ Cost Optimization: Real-time cost tracking and automatic optimization
  • โšก Rate Limiting: Sliding window rate limiting with key rotation
  • ๐Ÿ“ˆ Performance: Sub-millisecond routing with comprehensive monitoring
  • ๐Ÿ”’ Security: API key masking and secure storage
  • ๐Ÿ“Š Observability: Prometheus metrics, structured JSON logging

โ ๐Ÿค Contributing

We welcome contributions! Please see our Contributing Guidelinesโ  for details.

โ Development Setup
git clone https://github.com/your-org/coo-llm.git
cd coo-llm
go mod download
go build ./...
go test ./...
โ Key Areas for Contribution
  • ๐Ÿ”Œ New Providers: Add support for more LLM providers
  • โš–๏ธ Load Balancing: Improve routing algorithms
  • ๐Ÿ“Š Metrics: Add more observability features
  • ๐Ÿ”’ Security: Enhance security and authentication
  • ๐Ÿ“š Documentation: Improve docs and examples

โ ๐Ÿ“„ License

This project is licensed under the DIB License v1.0 - see the LICENSEโ  file for details.

โ ๐Ÿ™ Acknowledgments

  • OpenAI for the API specification that enables interoperability
  • Google & Anthropic for their excellent LLM APIs
  • The Go Community for outstanding tooling and libraries
  • LangChain for inspiring the integration examples
  • All Contributors who help make COO-LLM better

โ ๐Ÿ“ž Support & Community

โ ๐Ÿ† Key Highlights

  • ๐Ÿš€ Production Ready: Used in production with millions of requests
  • โšก High Performance: Sub-millisecond routing with Go's efficiency
  • ๐Ÿ”ง Easy Configuration: YAML-based config with environment variables
  • ๐Ÿ“Š Enterprise Observability: Prometheus metrics and structured logging
  • ๐Ÿ”„ Auto-Scaling: Horizontal scaling with Redis-backed state
  • ๐Ÿ’ฐ Cost Effective: Intelligent routing saves 20-50% on API costs

COO-LLM - The Intelligent LLM API Load Balancer ๐Ÿš€

Load balance your LLM API calls across multiple providers with OpenAI compatibility, real-time cost optimization, and enterprise-grade reliability.

Tag summary

Content type

Image

Digest

sha256:99e808b2fโ€ฆ

Size

46.9 MB

Last updated

12 months ago

docker pull khapu2906/coo-llm