Sign inSign up

himanshu806/ai-inference-platform

By himanshu806

•Updated 8 months ago

routes to best models by complexity. Cuts cost & latency.

Image
Machine learning & AI
Developer tools
0

478

himanshu806/ai-inference-platform repository overview

⁠AI Inference Platform

A high-performance AI inference platform built in Rust that intelligently routes queries to appropriate AI models (Claude, Gemini, ChatGPT, DeepSeek, Kimi) based on complexity classification, with caching and comprehensive monitoring capabilities.

Note

: This is a *super early release* (latest beta release version: **0.1.0**). It will be stable at **1.0.0** and ready for production mode. Drawbacks currently include: limited classification robustness, basic cache management, and early-stage telemetry.

More robust classification, improved cache manager, better telemetrics, and more configuration are coming. If you need a feature, mail me at: [email protected]⁠

⁠Features

⁠Core Capabilities
  • Vector-Based Query Classification: Uses embeddings and heuristics to classify queries as simple or complex
  • Intelligent Model Routing: Routes simple queries to smaller, faster models and complex queries to larger, more capable models
  • Multi-Level Caching: Redis-backed distributed caching with local in-memory cache for hot data
  • Semantic Cache Matching: Finds similar cached queries using cosine similarity on embeddings
  • Monitoring & Observability: Prometheus metrics, structured logging, and Grafana dashboards
⁠Architecture Highlights
┌─────────────┐     ┌─────────────────┐     ┌─────────────────┐
│   Client    │────▶│  Axum HTTP API  │────▶│  Query Classifier│
└─────────────┘     └─────────────────┘     └─────────────────┘
                                                        │
                       ┌────────────────────────────────┘
                       ▼
              ┌─────────────────┐
              │  Cache Manager  │
              │  (Redis + Local)│
              └─────────────────┘
                       │
              ┌────────┴────────┐
              ▼                 ▼
    ┌─────────────────┐ ┌─────────────────┐
    │  Simple Model   │ │  Complex Model  │
    └─────────────────┘ └─────────────────┘

⁠Quick Start

⁠Prerequisites
  • Rust 1.75+ (for building from source)
  • Docker & Docker Compose (for containerized deployment)
  • OpenAI API key (or other compatible API)
⁠Using Docker Compose
# Run standalone
docker run -p 8080:8080 himanshu806/ai-inference-platform:latest

# Run with full stack (Redis, Prometheus, Grafana)
curl -O https://raw.githubusercontent.com/yourusername/ai-inference-platform/main/docker-compose.yml
docker-compose up -d

⁠API Documentation

⁠Inference Endpoint

POST /api/v1/infer

Submit a query for AI inference. The platform will automatically classify the query complexity and route to the appropriate model.

⁠Request Body
{
  "query": "Explain quantum computing",
  "context": "optional context",
  "max_tokens": 500,
  "temperature": 0.7,
  "user_id": "user-123",
  "session_id": "session-456",
  "skip_cache": false
}
⁠Response
{
  "success": true,
  "data": {
    "response": "Quantum computing is a type of computing that uses quantum mechanics...",
    "model_used": "gpt-4",
    "complexity": "complex",
    "confidence": 0.85,
    "latency_ms": 1250,
    "cached": false,
    "tokens_used": {
      "prompt_tokens": 15,
      "completion_tokens": 150,
      "total_tokens": 165,
      "estimated_cost": 0.0099
    }
  },
  "error": null,
  "request_id": "uuid-here",
  "timestamp": "2024-01-15T10:30:00Z"
}
⁠Classification Endpoint

POST /api/v1/classify

Classify a query without performing inference.

⁠Request Body
{
  "query": "What is the weather today?"
}
⁠Response
{
  "success": true,
  "data": {
    "complexity": "simple",
    "confidence": 0.92,
    "scores": {
      "length_score": 0.1,
      "vocabulary_score": 0.2,
      "structure_score": 0.15,
      "semantic_score": 0.25,
      "keyword_score": 0.8,
      "final_score": 0.25
    }
  },
  "error": null,
  "request_id": "uuid-here",
  "timestamp": "2024-01-15T10:30:00Z"
}
⁠Health Check

GET /health

Returns service health status.

{
  "success": true,
  "data": {
    "status": "healthy",
    "version": "0.1.0",
    "uptime_seconds": 3600,
    "services": {
      "redis": true,
      "simple_model": true,
      "complex_model": true
    }
  },
  "error": null,
  "request_id": "uuid-here",
  "timestamp": "2024-01-15T10:30:00Z"
}
⁠Metrics

GET /metrics

Prometheus-compatible metrics endpoint.

⁠Statistics

GET /api/stats

Platform statistics including cache and router stats.

⁠Configuration

Configuration is managed through TOML files and environment variables.

⁠Query Classification

The platform uses a multi-factor approach to classify query complexity:

⁠Classification Factors
  1. Length Score (20%): Query length and word count
  2. Vocabulary Score (25%): Word diversity, average word length, technical terms
  3. Structure Score (25%): Sentence count, punctuation complexity, list indicators
  4. Semantic Score (30%): Embedding-based analysis (when available)
  5. Keyword Score (boost): Matches against simple/complex keyword lists
⁠Complexity Threshold

Queries with a final score >= complexity_threshold (default: 0.6) are classified as Complex and routed to the larger model. Others are classified as Simple and routed to the smaller model.

⁠Custom Keywords

You can configure custom keywords in the configuration:

[classifier]
simple_model_keywords = ["hello", "what is", "who is", "when", "where"]
complex_model_keywords = ["analyze", "compare", "code", "implement", "algorithm"]

⁠Caching

⁠Two-Level Cache Architecture
  1. Local Cache (L1): In-memory LRU cache using Moka

    • Fastest access
    • Limited by memory
    • Per-instance
  2. Distributed Cache (L2): Redis

    • Shared across instances
    • Persistent
    • Semantic matching support
⁠Semantic Cache Matching

The platform can find similar queries using cosine similarity on embeddings:

// Example: "What is machine learning?" matches "Explain machine learning"
// if similarity >= similarity_threshold (default: 0.95)
⁠Cache Configuration
[cache]
enable_local_cache = true
local_cache_size_mb = 100
similarity_threshold = 0.95
default_ttl_secs = 3600

⁠Monitoring

⁠Prometheus Metrics

The platform exposes the following metrics at /metrics:

MetricTypeDescription
inference_requests_totalCounterTotal inference requests
inference_errors_totalCounterTotal inference errors
inference_duration_secondsHistogramInference latency
cache_hits_totalCounterCache hit count
cache_misses_totalCounterCache miss count
active_requestsGaugeCurrently active requests
cache_sizeGaugeCurrent cache size
model_healthGaugeModel health status
⁠Grafana Dashboard

A pre-configured Grafana dashboard is included in monitoring/grafana/dashboards/.

⁠Accessing Monitoring

⁠Performance

⁠Benchmarks

Run benchmarks with:

cargo bench
⁠Benchmark Results (v0.1.0 Beta)

In our latest 100,000 request load test bypassing network LLM calls with a 5ms inline mock, the async routing platform achieved:

  • Raw Routing Throughput: 908.27 requests/second at 200 concurrency.
  • Reliability: 100% success rate (0 dropped packets out of 100k continuous requests).
  • Inherent Memory Footprint: ~38 MB peak resident memory under heavy sustained async load.
⁠Expected Performance (Live)
  • Inference Latency: 100ms-2000ms (depends on model: Claude, Gemini, ChatGPT, DeepSeek, Kimi)
  • Classification Latency: 1ms-50ms (depends on embedding model)
  • Cache Lookup: <1ms (local), 1-5ms (Redis)
⁠Optimization Tips
  1. Enable local cache for high-throughput scenarios
  2. Tune similarity_threshold based on your use case
  3. Use connection pooling for Redis
  4. Enable HTTP/2 for model API connections

⁠Deployment

⁠Kubernetes

Example deployment manifests:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: ai-inference-platform
spec:
  replicas: 3
  selector:
    matchLabels:
      app: ai-inference-platform
  template:
    metadata:
      labels:
        app: ai-inference-platform
    spec:
      containers:
      - name: app
        image: ai-inference-platform:latest
        ports:
        - containerPort: 8080
        env:
        - name: APP_CACHE__REDIS_URL
          value: "redis://redis-service:6379"
        - name: APP_MODELS__SIMPLE_MODEL__API_KEY
          valueFrom:
            secretKeyRef:
              name: api-keys
              key: openai
⁠Scaling

The platform is stateless and can be horizontally scaled:

# Scale to 5 instances
docker-compose up -d --scale ai-inference-platform=5

⁠Troubleshooting

⁠Common Issues

Redis Connection Failed

Failed to connect to Redis: Connection refused
  • Check Redis is running: docker-compose ps redis
  • Verify connection URL in config

Model API Errors

API error (401): Invalid authentication
  • Verify API keys are set correctly
  • Check API key has necessary permissions

High Memory Usage

  • Reduce local_cache_size_mb
  • Lower max_cached_queries
  • Enable cache eviction

⁠Acknowledgments

Tag summary

Content type

Image

Digest

sha256:087ac4f29…

Size

37.4 MB

Last updated

8 months ago

docker pull himanshu806/ai-inference-platform