routes to best models by complexity. Cuts cost & latency.
478
A high-performance AI inference platform built in Rust that intelligently routes queries to appropriate AI models (Claude, Gemini, ChatGPT, DeepSeek, Kimi) based on complexity classification, with caching and comprehensive monitoring capabilities.
Note
: This is a *super early release* (latest beta release version: **0.1.0**). It will be stable at **1.0.0** and ready for production mode. Drawbacks currently include: limited classification robustness, basic cache management, and early-stage telemetry.More robust classification, improved cache manager, better telemetrics, and more configuration are coming. If you need a feature, mail me at: [email protected]
┌─────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ Client │────▶│ Axum HTTP API │────▶│ Query Classifier│
└─────────────┘ └─────────────────┘ └─────────────────┘
│
┌────────────────────────────────┘
▼
┌─────────────────┐
│ Cache Manager │
│ (Redis + Local)│
└─────────────────┘
│
┌────────┴────────┐
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ Simple Model │ │ Complex Model │
└─────────────────┘ └─────────────────┘
# Run standalone
docker run -p 8080:8080 himanshu806/ai-inference-platform:latest
# Run with full stack (Redis, Prometheus, Grafana)
curl -O https://raw.githubusercontent.com/yourusername/ai-inference-platform/main/docker-compose.yml
docker-compose up -d
POST /api/v1/infer
Submit a query for AI inference. The platform will automatically classify the query complexity and route to the appropriate model.
{
"query": "Explain quantum computing",
"context": "optional context",
"max_tokens": 500,
"temperature": 0.7,
"user_id": "user-123",
"session_id": "session-456",
"skip_cache": false
}
{
"success": true,
"data": {
"response": "Quantum computing is a type of computing that uses quantum mechanics...",
"model_used": "gpt-4",
"complexity": "complex",
"confidence": 0.85,
"latency_ms": 1250,
"cached": false,
"tokens_used": {
"prompt_tokens": 15,
"completion_tokens": 150,
"total_tokens": 165,
"estimated_cost": 0.0099
}
},
"error": null,
"request_id": "uuid-here",
"timestamp": "2024-01-15T10:30:00Z"
}
POST /api/v1/classify
Classify a query without performing inference.
{
"query": "What is the weather today?"
}
{
"success": true,
"data": {
"complexity": "simple",
"confidence": 0.92,
"scores": {
"length_score": 0.1,
"vocabulary_score": 0.2,
"structure_score": 0.15,
"semantic_score": 0.25,
"keyword_score": 0.8,
"final_score": 0.25
}
},
"error": null,
"request_id": "uuid-here",
"timestamp": "2024-01-15T10:30:00Z"
}
GET /health
Returns service health status.
{
"success": true,
"data": {
"status": "healthy",
"version": "0.1.0",
"uptime_seconds": 3600,
"services": {
"redis": true,
"simple_model": true,
"complex_model": true
}
},
"error": null,
"request_id": "uuid-here",
"timestamp": "2024-01-15T10:30:00Z"
}
GET /metrics
Prometheus-compatible metrics endpoint.
GET /api/stats
Platform statistics including cache and router stats.
Configuration is managed through TOML files and environment variables.
The platform uses a multi-factor approach to classify query complexity:
Queries with a final score >= complexity_threshold (default: 0.6) are classified as Complex and routed to the larger model. Others are classified as Simple and routed to the smaller model.
You can configure custom keywords in the configuration:
[classifier]
simple_model_keywords = ["hello", "what is", "who is", "when", "where"]
complex_model_keywords = ["analyze", "compare", "code", "implement", "algorithm"]
Local Cache (L1): In-memory LRU cache using Moka
Distributed Cache (L2): Redis
The platform can find similar queries using cosine similarity on embeddings:
// Example: "What is machine learning?" matches "Explain machine learning"
// if similarity >= similarity_threshold (default: 0.95)
[cache]
enable_local_cache = true
local_cache_size_mb = 100
similarity_threshold = 0.95
default_ttl_secs = 3600
The platform exposes the following metrics at /metrics:
| Metric | Type | Description |
|---|---|---|
inference_requests_total | Counter | Total inference requests |
inference_errors_total | Counter | Total inference errors |
inference_duration_seconds | Histogram | Inference latency |
cache_hits_total | Counter | Cache hit count |
cache_misses_total | Counter | Cache miss count |
active_requests | Gauge | Currently active requests |
cache_size | Gauge | Current cache size |
model_health | Gauge | Model health status |
A pre-configured Grafana dashboard is included in monitoring/grafana/dashboards/.
Run benchmarks with:
cargo bench
In our latest 100,000 request load test bypassing network LLM calls with a 5ms inline mock, the async routing platform achieved:
Example deployment manifests:
apiVersion: apps/v1
kind: Deployment
metadata:
name: ai-inference-platform
spec:
replicas: 3
selector:
matchLabels:
app: ai-inference-platform
template:
metadata:
labels:
app: ai-inference-platform
spec:
containers:
- name: app
image: ai-inference-platform:latest
ports:
- containerPort: 8080
env:
- name: APP_CACHE__REDIS_URL
value: "redis://redis-service:6379"
- name: APP_MODELS__SIMPLE_MODEL__API_KEY
valueFrom:
secretKeyRef:
name: api-keys
key: openai
The platform is stateless and can be horizontally scaled:
# Scale to 5 instances
docker-compose up -d --scale ai-inference-platform=5
Redis Connection Failed
Failed to connect to Redis: Connection refused
docker-compose ps redisModel API Errors
API error (401): Invalid authentication
High Memory Usage
local_cache_size_mbmax_cached_queriesContent type
Image
Digest
sha256:087ac4f29…
Size
37.4 MB
Last updated
8 months ago
docker pull himanshu806/ai-inference-platform