OpenAI-compatible router that routes text and vision requests to the best on-prem LLM
6.3K
This service is an OpenAI-compatible API router designed to intelligently dispatch requests to different vLLM backends based on request content and model capabilities.
It enables a single entry point (for example OpenWebUI) to transparently consume multiple on-premise LLMs, each optimized for a specific task such as text generation or image understanding, without requiring any client-side changes.
The router solves a common on-premise limitation:
no single model excels at everything.
For example:
This service automatically routes:
The client continues to use the standard OpenAI Chat Completions API, unaware of the internal routing logic.
/v1/chat/completions and /v1/completions endpointsmodel field when neededThe router operates in three modes:
If the request specifies the configured text or image model name, the request is routed accordingly.
model=auto)The router can rewrite the model field before forwarding the request to match the exact model name expected by the vLLM backend.
URL_LLM_TXT
Base URL of the text-only vLLM backend (with or without /v1)
APIKEY_LLM_TXT (optional)
API key injected as Authorization: Bearer … when forwarding requests
MODEL_NAME_LLM_TXT
Model name exposed to clients (e.g. oss-20b)
UPSTREAM_MODEL_NAME_LLM_TXT (optional)
Actual model name expected by the vLLM backend
URL_LLM_IMG
Base URL of the vision-capable vLLM backend
APIKEY_LLM_IMG (optional)
API key injected when forwarding requests
MODEL_NAME_LLM_IMG
Model name exposed to clients (e.g. gemma-vision)
UPSTREAM_MODEL_NAME_LLM_IMG (optional)
Actual upstream model name used by the vision backend
EXPOSE_AUTO_MODEL (default: true)
Exposes the virtual auto model in /v1/models
MAX_IMAGE_URL_CHARS (default: 20000000)
Safety limit to prevent oversized data URLs
This router is designed to work seamlessly with:
No client-side modification is required.
Detailed architecture explanations and real-world use cases are available :
This service acts as a capability-aware LLM gateway, enabling:
It allows organizations to combine the strengths of multiple models while preserving a single, stable OpenAI-compatible API surface.
This section provides a complete example to run the router with Docker Compose, including all supported environment variables.
Notes:
- The router listens on port 8000 inside the container.
URL_LLM_TXT/URL_LLM_IMGmay be provided with or without/v1(the router normalizes them).- API keys are optional. If the client already sends an
Authorizationheader, it is forwarded as-is.
version: "3.9"
services:
modelsrouter:
image: suntux57420/modelsrouter:latest
container_name: modelsrouter
restart: unless-stopped
ports:
- "8000:8000"
environment:
# ---- Text backend (required) ----
URL_LLM_TXT: "http://oss-vllm:8000/v1"
MODEL_NAME_LLM_TXT: "oss-20b"
APIKEY_LLM_TXT: "" # optional
# Optional: rewrite "model" before forwarding to the TXT backend
# Example: client uses "oss-20b" but vLLM expects "oss-20b-instruct"
UPSTREAM_MODEL_NAME_LLM_TXT: "" # optional
# ---- Image/Vision backend (required) ----
URL_LLM_IMG: "http://gemma-vllm:8000/v1"
MODEL_NAME_LLM_IMG: "gemma-vision"
APIKEY_LLM_IMG: "" # optional
# Optional: rewrite "model" before forwarding to the IMG backend
UPSTREAM_MODEL_NAME_LLM_IMG: "" # optional
# ---- Optional behavior ----
EXPOSE_AUTO_MODEL: "true"
MAX_IMAGE_URL_CHARS: "20000000"
# Optional: attach to an explicit network shared with vLLM backends
# networks:
# - llm_net
# Optional examples of backends (replace by your own vLLM services)
# oss-vllm:
# image: vllm/vllm-openai:latest
# ...
# gemma-vllm:
# image: vllm/vllm-openai:latest
# ...
# networks:
# llm_net:
# driver: bridge
docker compose up -d
docker compose logs -f modelsrouter
docker compose down
curl -s http://127.0.0.1:8000/v1/models
This section provides a ready-to-apply Kubernetes YAML including:
DeploymentServiceSecret for API keysIt also includes all supported environment variables.
apiVersion: apps/v1
kind: Deployment
metadata:
name: modelsrouter
labels:
app: modelsrouter
spec:
replicas: 1
selector:
matchLabels:
app: modelsrouter
template:
metadata:
labels:
app: modelsrouter
spec:
containers:
- name: modelsrouter
image: suntux57420/modelsrouter:latest
imagePullPolicy: IfNotPresent
ports:
- containerPort: 8000
env:
# ---- Text backend ----
- name: URL_LLM_TXT
value: "http://oss-vllm.default.svc.cluster.local:8000/v1"
- name: MODEL_NAME_LLM_TXT
value: "oss-20b"
- name: APIKEY_LLM_TXT
value: "" # optional
- name: UPSTREAM_MODEL_NAME_LLM_TXT
value: "" # optional
# ---- Image/Vision backend ----
- name: URL_LLM_IMG
value: "http://gemma-vllm.default.svc.cluster.local:8000/v1"
- name: MODEL_NAME_LLM_IMG
value: "gemma-vision"
- name: APIKEY_LLM_IMG
value: "" # optional
- name: UPSTREAM_MODEL_NAME_LLM_IMG
value: "" # optional
# ---- Optional behavior ----
- name: EXPOSE_AUTO_MODEL
value: "true"
- name: MAX_IMAGE_URL_CHARS
value: "20000000"
---
apiVersion: v1
kind: Service
metadata:
name: modelsrouter
labels:
app: modelsrouter
spec:
selector:
app: modelsrouter
ports:
- name: http
port: 8000
targetPort: 8000
type: ClusterIP
Apply it:
kubectl apply -f modelsrouter.yaml
kubectl get pod -l app=modelsrouter
kubectl get svc modelsrouter
Create a secret:
apiVersion: v1
kind: Secret
metadata:
name: modelsrouter-secrets
type: Opaque
stringData:
APIKEY_LLM_TXT: ""
APIKEY_LLM_IMG: ""
Then reference it in the Deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
name: modelsrouter
labels:
app: modelsrouter
spec:
replicas: 1
selector:
matchLabels:
app: modelsrouter
template:
metadata:
labels:
app: modelsrouter
spec:
containers:
- name: modelsrouter
image: suntux57420/modelsrouter:latest
ports:
- containerPort: 8000
env:
# ---- Text backend ----
- name: URL_LLM_TXT
value: "http://oss-vllm.default.svc.cluster.local:8000/v1"
- name: MODEL_NAME_LLM_TXT
value: "oss-20b"
- name: UPSTREAM_MODEL_NAME_LLM_TXT
value: "" # optional
# Injected from Secret
- name: APIKEY_LLM_TXT
valueFrom:
secretKeyRef:
name: modelsrouter-secrets
key: APIKEY_LLM_TXT
# ---- Image/Vision backend ----
- name: URL_LLM_IMG
value: "http://gemma-vllm.default.svc.cluster.local:8000/v1"
- name: MODEL_NAME_LLM_IMG
value: "gemma-vision"
- name: UPSTREAM_MODEL_NAME_LLM_IMG
value: "" # optional
# Injected from Secret
- name: APIKEY_LLM_IMG
valueFrom:
secretKeyRef:
name: modelsrouter-secrets
key: APIKEY_LLM_IMG
# ---- Optional behavior ----
- name: EXPOSE_AUTO_MODEL
value: "true"
- name: MAX_IMAGE_URL_CHARS
value: "20000000"
---
apiVersion: v1
kind: Service
metadata:
name: modelsrouter
labels:
app: modelsrouter
spec:
selector:
app: modelsrouter
ports:
- name: http
port: 8000
targetPort: 8000
type: ClusterIP
Apply it:
kubectl apply -f modelsrouter-secret.yaml
kubectl apply -f modelsrouter.yaml
kubectl run -it --rm curl --image=curlimages/curl --restart=Never -- curl -s http://modelsrouter:8000/v1/models
When configuring an OpenAI-compatible endpoint in OpenWebUI:
http://<router-host>:8000/v1auto (or select the explicit exposed models)Content type
Image
Digest
sha256:1db9e89c1…
Size
56.4 MB
Last updated
8 months ago
docker pull suntux57420/modelsrouter