Sign inSign up

schedion/llmrouter

By schedion

•Updated about 1 year ago

llmrouter multi-provider router

Image
Machine learning & AI
0

988

schedion/llmrouter repository overview

⁠llmrouter

llmrouter is a programmable OpenAI-compatible router for Large Language Models. It lets you combine multiple providers, fall back gracefully, and cache responses without changing upstream client code.

⁠Supported Providers (out of the box)

⁠Quick Start

# install dependencies
pip install -r requirements.txt
# optional: install semantic cache extras (pulls in PyTorch)
pip install -r requirements-semantic.txt

# export provider credentials
export PROVIDER_KEY_GROQ=...
export PROVIDER_KEY_OPENROUTER=...
export PROVIDER_KEY_NVIDIA_NIM=...
export PROVIDER_KEY_HUGGINGFACE=...

# point at the published catalog (or your local one)
export LLMROUTER_MODEL_INDEX_URL="https://raw.githubusercontent.com/schedion/llmrouter/refs/heads/main/generated/model_index.json"

uvicorn app.main:app --reload

Query the API:

# list available canonical models
curl http://localhost:8000/v1/models

# run a chat completion
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-oss-20b",
        "messages": [
          {"role": "system", "content": "You are a concise assistant."},
          {"role": "user", "content": "hello"}
        ]
      }'

Semantic caching controls:

  • enable/disable globally with LLMROUTER_SEMANTIC_CACHE_ENABLED (false by default)
  • override per request using headers:
    • X-LLMRouter-Semantic-Cache: on|off
    • X-LLMRouter-Semantic-Threshold: 0.9
  • Looking to slim the Docker image? Today we rely on sentence-transformers, which pulls in PyTorch. We may switch to an ONNXRuntime-backed embedding loader in the future; if you explore that path, keep the MiniLM model and tokenizer files together and use onnxruntime for inference.

Exact-match caching uses LLMROUTER_CACHE_TTL (seconds, 0 disables).

⁠Configuration

  • Canonical model mappings live in config/model_catalog_seed.yaml.
  • Regenerate the catalog (generated/model_index.json) with:
    ./scripts/build_free_model_catalog.py \
      --providers groq,openrouter,nvidia_nim,huggingface \
      --output generated/model_index.json --pretty
    
  • Skip providers you don’t use by editing the seed or passing a smaller --providers list.
  • Provider credentials expected by default: PROVIDER_KEY_GROQ, PROVIDER_KEY_OPENROUTER, PROVIDER_KEY_NVIDIA_NIM, PROVIDER_KEY_HUGGINGFACE.

⁠Docker

A sample Compose file (docker-compose.sample.yml) mounts named volumes for config, generated data, and the Hugging Face cache. Update the image name and credentials, then run:

docker compose -f docker-compose.sample.yml up -d

Published images:

  • schedion/llmrouter:slim – no semantic cache dependencies, smallest footprint (multi-arch: linux/amd64, linux/arm64, linux/arm/v7).
  • schedion/llmrouter:latest – includes semantic cache extras (published for linux/amd64 only because PyTorch wheels aren’t available for our other targets).

For the arm/v7 build we pull Python wheels from piwheels.org⁠; if you build locally on a Raspberry Pi, keep the same extra index (PIP_EXTRA_INDEX_URL=https://www.piwheels.org/simple) to avoid compiling NumPy from source.

⁠Observability
  • Prometheus metrics are exposed at /metrics (suitable for scraping).
  • A lightweight HTML dashboard at /metrics/dashboard summarises API traffic and upstream provider status/latency for quick local inspection.

⁠Publishing to Docker Hub

⁠Development

python -m compileall app scripts  # quick syntax check
pytest                             # when tests are added

⁠License

The Unlicense⁠

Tag summary

Content type

Image

Digest

sha256:90e677a69…

Size

3.9 GB

Last updated

about 1 year ago

docker pull schedion/llmrouter