Sign inSign up

suntux57420/modelsrouter

By suntux57420

•Updated 8 months ago

OpenAI-compatible router that routes text and vision requests to the best on-prem LLM

Image
Integration & delivery
API management
Machine learning & AI
0

6.3K

suntux57420/modelsrouter repository overview

⁠OpenAI-Compatible LLM Model Router

⁠Overview

This service is an OpenAI-compatible API router designed to intelligently dispatch requests to different vLLM backends based on request content and model capabilities.

It enables a single entry point (for example OpenWebUI) to transparently consume multiple on-premise LLMs, each optimized for a specific task such as text generation or image understanding, without requiring any client-side changes.

⁠Purpose

The router solves a common on-premise limitation:
no single model excels at everything.

For example:

  • A vision-capable model (e.g. Gemma 3 Vision) can analyze images accurately but may produce weaker textual reasoning.
  • A text-optimized model (e.g. OSS-20B) can generate high-quality answers but does not support image inputs.

This service automatically routes:

  • requests containing images → to the vision-capable model
  • text-only requests → to the text-optimized model

The client continues to use the standard OpenAI Chat Completions API, unaware of the internal routing logic.

⁠Key Features

  • OpenAI-compatible /v1/chat/completions and /v1/completions endpoints
  • Automatic detection of image inputs (multimodal or data URLs)
  • Explicit model forcing via model field when needed
  • Transparent SSE streaming relay
  • Optional API key injection per backend
  • Fully configurable via environment variables
  • Container-ready (Docker / Kubernetes friendly)

⁠Model Selection Logic

The router operates in three modes:

⁠1. Explicit model selection

If the request specifies the configured text or image model name, the request is routed accordingly.

⁠2. Automatic routing (model=auto)
  • If an image is detected in the request, it is routed to the image model.
  • Otherwise, it is routed to the text model.
⁠3. Upstream model rewriting (optional)

The router can rewrite the model field before forwarding the request to match the exact model name expected by the vLLM backend.

⁠Environment Variables

⁠Text LLM backend
  • URL_LLM_TXT
    Base URL of the text-only vLLM backend (with or without /v1)

  • APIKEY_LLM_TXT (optional)
    API key injected as Authorization: Bearer … when forwarding requests

  • MODEL_NAME_LLM_TXT
    Model name exposed to clients (e.g. oss-20b)

  • UPSTREAM_MODEL_NAME_LLM_TXT (optional)
    Actual model name expected by the vLLM backend

⁠Image / Vision LLM backend
  • URL_LLM_IMG
    Base URL of the vision-capable vLLM backend

  • APIKEY_LLM_IMG (optional)
    API key injected when forwarding requests

  • MODEL_NAME_LLM_IMG
    Model name exposed to clients (e.g. gemma-vision)

  • UPSTREAM_MODEL_NAME_LLM_IMG (optional)
    Actual upstream model name used by the vision backend

⁠Optional behavior
  • EXPOSE_AUTO_MODEL (default: true)
    Exposes the virtual auto model in /v1/models

  • MAX_IMAGE_URL_CHARS (default: 20000000)
    Safety limit to prevent oversized data URLs

⁠Compatible Clients

This router is designed to work seamlessly with:

  • OpenWebUI
  • Custom OpenAI-compatible clients
  • Internal RAG gateways
  • Any application consuming the OpenAI Chat Completions API

No client-side modification is required.

⁠Documentation & Use Cases

Detailed architecture explanations and real-world use cases are available :

⁠Summary

This service acts as a capability-aware LLM gateway, enabling:

  • optimal model usage
  • clean separation of concerns
  • scalable on-premise AI architectures

It allows organizations to combine the strengths of multiple models while preserving a single, stable OpenAI-compatible API surface.

⁠Running with Docker (Compose)

This section provides a complete example to run the router with Docker Compose, including all supported environment variables.

⁠Example docker-compose.yml

Notes:

  • The router listens on port 8000 inside the container.
  • URL_LLM_TXT / URL_LLM_IMG may be provided with or without /v1 (the router normalizes them).
  • API keys are optional. If the client already sends an Authorization header, it is forwarded as-is.
version: "3.9"

services:
  modelsrouter:
    image: suntux57420/modelsrouter:latest
    container_name: modelsrouter
    restart: unless-stopped
    ports:
      - "8000:8000"
    environment:
      # ---- Text backend (required) ----
      URL_LLM_TXT: "http://oss-vllm:8000/v1"
      MODEL_NAME_LLM_TXT: "oss-20b"
      APIKEY_LLM_TXT: ""  # optional

      # Optional: rewrite "model" before forwarding to the TXT backend
      # Example: client uses "oss-20b" but vLLM expects "oss-20b-instruct"
      UPSTREAM_MODEL_NAME_LLM_TXT: ""  # optional

      # ---- Image/Vision backend (required) ----
      URL_LLM_IMG: "http://gemma-vllm:8000/v1"
      MODEL_NAME_LLM_IMG: "gemma-vision"
      APIKEY_LLM_IMG: ""  # optional

      # Optional: rewrite "model" before forwarding to the IMG backend
      UPSTREAM_MODEL_NAME_LLM_IMG: ""  # optional

      # ---- Optional behavior ----
      EXPOSE_AUTO_MODEL: "true"
      MAX_IMAGE_URL_CHARS: "20000000"

    # Optional: attach to an explicit network shared with vLLM backends
    # networks:
    #   - llm_net

  # Optional examples of backends (replace by your own vLLM services)
  # oss-vllm:
  #   image: vllm/vllm-openai:latest
  #   ...

  # gemma-vllm:
  #   image: vllm/vllm-openai:latest
  #   ...

# networks:
#   llm_net:
#     driver: bridge

⁠Start / Stop

docker compose up -d
docker compose logs -f modelsrouter
docker compose down

⁠Quick test

curl -s http://127.0.0.1:8000/v1/models

⁠Running on Kubernetes

This section provides a ready-to-apply Kubernetes YAML including:

  • a Deployment
  • a Service
  • an optional Secret for API keys

It also includes all supported environment variables.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: modelsrouter
  labels:
    app: modelsrouter
spec:
  replicas: 1
  selector:
    matchLabels:
      app: modelsrouter
  template:
    metadata:
      labels:
        app: modelsrouter
    spec:
      containers:
        - name: modelsrouter
          image: suntux57420/modelsrouter:latest
          imagePullPolicy: IfNotPresent
          ports:
            - containerPort: 8000
          env:
            # ---- Text backend ----
            - name: URL_LLM_TXT
              value: "http://oss-vllm.default.svc.cluster.local:8000/v1"
            - name: MODEL_NAME_LLM_TXT
              value: "oss-20b"
            - name: APIKEY_LLM_TXT
              value: ""  # optional
            - name: UPSTREAM_MODEL_NAME_LLM_TXT
              value: ""  # optional

            # ---- Image/Vision backend ----
            - name: URL_LLM_IMG
              value: "http://gemma-vllm.default.svc.cluster.local:8000/v1"
            - name: MODEL_NAME_LLM_IMG
              value: "gemma-vision"
            - name: APIKEY_LLM_IMG
              value: ""  # optional
            - name: UPSTREAM_MODEL_NAME_LLM_IMG
              value: ""  # optional

            # ---- Optional behavior ----
            - name: EXPOSE_AUTO_MODEL
              value: "true"
            - name: MAX_IMAGE_URL_CHARS
              value: "20000000"
---
apiVersion: v1
kind: Service
metadata:
  name: modelsrouter
  labels:
    app: modelsrouter
spec:
  selector:
    app: modelsrouter
  ports:
    - name: http
      port: 8000
      targetPort: 8000
  type: ClusterIP

Apply it:

kubectl apply -f modelsrouter.yaml
kubectl get pod -l app=modelsrouter
kubectl get svc modelsrouter

Create a secret:

apiVersion: v1
kind: Secret
metadata:
  name: modelsrouter-secrets
type: Opaque
stringData:
  APIKEY_LLM_TXT: ""
  APIKEY_LLM_IMG: ""

Then reference it in the Deployment:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: modelsrouter
  labels:
    app: modelsrouter
spec:
  replicas: 1
  selector:
    matchLabels:
      app: modelsrouter
  template:
    metadata:
      labels:
        app: modelsrouter
    spec:
      containers:
        - name: modelsrouter
          image: suntux57420/modelsrouter:latest
          ports:
            - containerPort: 8000
          env:
            # ---- Text backend ----
            - name: URL_LLM_TXT
              value: "http://oss-vllm.default.svc.cluster.local:8000/v1"
            - name: MODEL_NAME_LLM_TXT
              value: "oss-20b"
            - name: UPSTREAM_MODEL_NAME_LLM_TXT
              value: ""  # optional

            # Injected from Secret
            - name: APIKEY_LLM_TXT
              valueFrom:
                secretKeyRef:
                  name: modelsrouter-secrets
                  key: APIKEY_LLM_TXT

            # ---- Image/Vision backend ----
            - name: URL_LLM_IMG
              value: "http://gemma-vllm.default.svc.cluster.local:8000/v1"
            - name: MODEL_NAME_LLM_IMG
              value: "gemma-vision"
            - name: UPSTREAM_MODEL_NAME_LLM_IMG
              value: ""  # optional

            # Injected from Secret
            - name: APIKEY_LLM_IMG
              valueFrom:
                secretKeyRef:
                  name: modelsrouter-secrets
                  key: APIKEY_LLM_IMG

            # ---- Optional behavior ----
            - name: EXPOSE_AUTO_MODEL
              value: "true"
            - name: MAX_IMAGE_URL_CHARS
              value: "20000000"
---
apiVersion: v1
kind: Service
metadata:
  name: modelsrouter
  labels:
    app: modelsrouter
spec:
  selector:
    app: modelsrouter
  ports:
    - name: http
      port: 8000
      targetPort: 8000
  type: ClusterIP

Apply it:

kubectl apply -f modelsrouter-secret.yaml
kubectl apply -f modelsrouter.yaml

⁠Quick test inside the cluster

kubectl run -it --rm curl --image=curlimages/curl --restart=Never --   curl -s http://modelsrouter:8000/v1/models

⁠Common Integration Tip (OpenWebUI)

When configuring an OpenAI-compatible endpoint in OpenWebUI:

  • Base URL: http://<router-host>:8000/v1
  • Model: auto (or select the explicit exposed models)
  • API key: optional (depending on your router exposure policy)

Tag summary

Content type

Image

Digest

sha256:1db9e89c1…

Size

56.4 MB

Last updated

8 months ago

docker pull suntux57420/modelsrouter