Multi-Provider AI Gateway - Model Autodiscovery, Failover & more - "Because we have LiteLLM at home"
10K+
github.com/hugalafutro/model-hotel
"Because we have LiteLLM at home"
Multi-Provider AI Gateway
AI-Assisted Project Disclaimer: Human judgment applied at every stage, particularly around architectural decisions, UX flows, and quality control.
A single OpenAI-compatible endpoint in front of all your LLM providers, cloud or self-hosted. Add the same model from any number of providers and a failover group forms around it automatically: requests go to the providers in the order you set, and when one runs out of quota or goes down, the next one answers, with no change on the client side. A provider whose quota is spent is skipped until it resets, then rejoins on its own. Models are auto-discovered the moment you add a provider and, if you want, on a schedule. No prompt data is ever stored. Full feature tour, screenshots, and the security/auth breakdown live on GitHub.
Live demo: poke around a real instance at mh.site19.ddns.net - rebuilds fresh every 30 minutes.
Two ways in: clone and build from source (below), or skip the clone and run the published image from two files (see Deploy without Git below).
git clone https://github.com/hugalafutro/model-hotel.git
cd model-hotel
cp .env.example .env
nano .env # set a strong MASTER_KEY and POSTGRES_PASSWORD; change HOST_PORT if 8081 is taken
docker compose up --build -d
For local development, layer the compose.dev.yml override instead. It mounts the Docker socket, turns on DEBUG_LOG, and allows embedding, so use it only in a trusted environment:
# Development only:
docker compose -f docker-compose.yml -f compose.dev.yml up --build -d
To run a prebuilt image instead of building from source, edit docker-compose.yml: comment out the build: block and uncomment one of the image: lines.
On first run the admin token is printed once, in a boxed ADMIN TOKEN block in the startup banner, and never shown again:
docker compose logs app
If you lose it, delete .data/admin-token and restart to generate a new one. The ADMIN_TOKEN environment variable seeds the token on first boot only: once .data/admin-token exists the file wins and the variable is ignored.
Open http://localhost:8081 (or the HOST_PORT you set), log in with that token, add your first provider, and start proxying. To stop, update or remove the stack, see Stop, Update, Remove below.
No git clone needed, and no build. Create two files and go:
1. Create .env with your secrets:
# Generate strong secrets:
# MASTER_KEY: openssl rand -base64 32
# POSTGRES_PASSWORD: openssl rand -hex 16
# ADMIN_TOKEN: openssl rand -hex 16 (optional; auto-generated if empty)
MASTER_KEY=<your-master-key>
POSTGRES_PASSWORD=<your-postgres-password>
ADMIN_TOKEN=
# Optional: host port for the dashboard and API; change it if 8081 is taken
# HOST_PORT=8081
# Optional: WebAuthn/FIDO2 passkey login (only WEBAUTHN_RP_ID is required)
# WEBAUTHN_RP_ID=your-domain.com
# WEBAUTHN_RP_ORIGINS=https://your-domain.com
2. Create docker-compose.yml:
name: model-hotel
services:
app:
# Build from source (default):
build:
context: .
args:
VERSION: ${VERSION:-dev}
COMMIT: ${COMMIT:-unknown}
# Prebuilt images (uncomment 1 image according to registry preference, comment out build above):
# image: ghcr.io/hugalafutro/model-hotel:latest
# image: hugalafutro/model-hotel:latest
labels:
app.group: model-hotel
ports:
- "${HOST_PORT:-8081}:8080"
environment:
- MASTER_KEY=${MASTER_KEY:?MASTER_KEY must be set in .env}
- POSTGRES_USER=${POSTGRES_USER:-modelhotel}
- POSTGRES_PASSWORD=${POSTGRES_PASSWORD:?POSTGRES_PASSWORD must be set in .env}
- POSTGRES_HOST=db
- POSTGRES_DB=${POSTGRES_DB:-modelhotel}
- ADMIN_TOKEN=${ADMIN_TOKEN:-}
- ALLOW_HTTP_PROVIDERS=false
- ALLOW_EMBED=false
- DATA_DIR=/data
- RATE_LIMIT_ENABLED=true
- DEBUG_LOG=false
- CORS_ORIGINS=http://localhost:5173,http://localhost:${HOST_PORT:-8081}
- WEBAUTHN_RP_ID=${WEBAUTHN_RP_ID:-}
- WEBAUTHN_RP_ORIGINS=${WEBAUTHN_RP_ORIGINS:-}
- ALLOWED_PROVIDER_HOSTS=
- TRUSTED_PROXIES=
- KNOWN_PROXIES=
volumes:
- ./.data:/data
# Docker socket (disabled by default for security).
# Enable to show container-level stats in the sidebar (CPU, memory per container).
# ⚠️ Granting Docker socket access allows the container to control the Docker daemon.
# Only enable if you trust the deployment environment.
# - /var/run/docker.sock:/var/run/docker.sock:ro
restart: unless-stopped
# Model Hotel winds down in stages on SIGTERM. Worst case, in order:
# 10s HTTP drain (open SSE tabs and proxied streams are ended first, so
# this is usually quick) + 35s background join (the 30s ceiling of the
# scheduled-disable sweep, which deliberately finishes the statement it
# has already started, plus a 5s margin; the retention and stale-log
# sweeps have no ceiling and the join cancels them instead of waiting)
# + 10s audit drain (one record's 5s insert plus the 5s retention prune
# it piggybacks) + 5s app-log writer stop + 5s OTLP flush = 65s. The
# closes around them (the event bus, the proxy handler, discovery, the
# docker client, the rate limiters and the database pool) carry no budget
# of their own, so this is a ceiling with headroom over the 65s, not the
# sum. Docker's default grace is 10s, which would SIGKILL partway through
# the drain and take the audit rows and the last log lines with it.
stop_grace_period: 75s
depends_on:
db:
condition: service_healthy
db:
image: postgres:16-alpine
labels:
app.group: model-hotel
command: ["postgres", "-c", "log_min_error_statement=panic", "-c", "log_min_messages=error", "-c", "log_checkpoints=off"]
environment:
- POSTGRES_USER=${POSTGRES_USER:-modelhotel}
- POSTGRES_PASSWORD=${POSTGRES_PASSWORD:?POSTGRES_PASSWORD must be set in .env}
- POSTGRES_DB=${POSTGRES_DB:-modelhotel}
volumes:
- ./.data/pgdata:/var/lib/postgresql/data
restart: unless-stopped
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ${POSTGRES_USER:-modelhotel}"]
interval: 5s
timeout: 5s
retries: 5
# Optional: outbound alerting via Apprise. Uncomment to run a stateless
# apprise-api container, then in Settings → Alerts switch alerting on and press
# "Set up alerts": the wizard checks http://apprise:8000, builds the destination
# URL for you (ntfy, Telegram, Discord, email, or a raw Apprise URL), tests it,
# and saves only at Finish. The same fields sit under "Manual configuration (advanced)" if you
# would rather paste tgram://<bot_token>/<chat_id> yourself. Model Hotel POSTs
# event summaries here and Apprise fans them out to your service. No request
# content is ever sent.
# apprise:
# image: caronc/apprise:latest
# labels:
# app.group: model-hotel
# restart: unless-stopped
# # Not exposed to the host: only Model Hotel needs to reach it.
# expose:
# - "8000"
3. Switch to the prebuilt image. The file above builds from source, which needs the repository next to it, so in your copy comment out the build: block (the build: line and the four lines under it) and uncomment one of the two image: lines (GHCR or Docker Hub).
4. Deploy:
docker compose up -d
Then read the admin token from docker compose logs app and open http://localhost:8081 (or the HOST_PORT you set). The file sets name: model-hotel, so a second stack on the same host needs a different name: (or -p <other> on every compose command) and a different HOST_PORT.
Note: The compose above is the production file; see Quick Start above for the development override.
WEBAUTHN_RP_IDenables passkey login (empty to disable);TRUSTED_PROXIEStrusts inboundX-Forwarded-Forheaders from reverse proxies;KNOWN_PROXIESallows outbound connections to internal LLM servers on private networks (bypasses SSRF protection). See the Configuration wiki for every variable.
Note: The app only sees the variables listed under its
environment:key;.envjust fills their${...}placeholders. To use any other variable (for exampleCOOKIE_SECURE,METRICS_TOKENorLOG_FORMAT), add it to that list, e.g.- COOKIE_SECURE=${COOKIE_SECURE:-always}.COOKIE_SECUREsets theSecureattribute on the dashboard login cookies:always(the default) sends them only over HTTPS or tohttp://localhost, so logging in over plain HTTP from another machine (e.g.http://192.168.1.10:8081) fails until you setauto(follows the request: TLS orX-Forwarded-Proto: https) ornever(plain-HTTP LAN).
docker compose down stops the stack and keeps your data. Both services use bind mounts under ./.data (PostgreSQL in ./.data/pgdata). There are no named volumes, so down -v removes nothing more.
To update a two-file deployment, run docker compose pull && docker compose up -d. Compose changes do not reach you on their own: diff the block in Deploy without Git against your file now and then, merge what changed by hand, keeping every local edit (the step 3 image switch, added environment: entries, an uncommented socket mount or apprise service), then run the same two commands. A clone that builds from source runs git pull && docker compose pull --ignore-buildable && docker compose up --build -d (the extra pull refreshes the PostgreSQL image, which up --build leaves alone). A clone switched to a prebuilt image has a local edit in docker-compose.yml, so run git stash && git pull && git stash pop; if the pop reports a conflict, remove the conflict markers in docker-compose.yml, keeping your image: line and upstream's other changes, then run git restore --staged docker-compose.yml && git stash drop. Finish with docker compose pull && docker compose up -d.
To remove everything, run docker compose down --rmi all (containers, network and the images the services use) and delete ./.data. ./.data/pgdata belongs to PostgreSQL (uid 70), so this needs sudo rm -rf .data; the rest of .data belongs to uid 1000, which usually matches your host user. .env and docker-compose.yml are yours to delete.
A single instance keeps its caches and rate limiters in memory, so to survive a host failure you run several instances behind one client endpoint. A Front Desk control plane holds the fleet roster and replicates config to every member, and Traefik load-balances them with health checks and automatic failover; members share one MASTER_KEY so encrypted provider keys port across the fleet. Full runbook in the High Availability guide.
Bellhop, the Android companion app, pairs with Front Desk for a pocket view of fleet health, traffic, provider quota badges, and events, plus operator controls behind a biometric prompt and a home-screen widget. See the Bellhop guide. APK download: (signed; Obtainium-compatible).
# List available models
curl http://localhost:8081/v1/models \
-H "Authorization: Bearer $VIRTUAL_KEY"
# Chat completion (hotel/ routing for automatic failover across providers)
curl -X POST http://localhost:8081/v1/chat/completions \
-H "Authorization: Bearer $VIRTUAL_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "hotel/glm-4.6", "messages": [{"role": "user", "content": "Hello!"}]}'
# Speech-to-text (multimodal endpoints share the same provider/model and hotel/ routing)
curl -X POST http://localhost:8081/v1/audio/transcriptions \
-H "Authorization: Bearer $VIRTUAL_KEY" \
-F model="OpenAI/whisper-1" -F [email protected]
The proxy also serves /v1/embeddings, /v1/rerank (Cohere's per-search billing is priced into spend and budgets), /v1/images/generations|edits|variations, /v1/audio/speech (TTS; Gemini TTS models are served through the native route and answer as wav or pcm), /v1/audio/translations, a native Anthropic POST /v1/messages (for Claude Code / anthropic SDKs, with cross-provider failover), and the OpenAI Responses API POST /v1/responses (for Codex CLI and Responses-only SDK clients; forwarded verbatim to OpenAI, translated for every other provider) - all with the same routing, failover and virtual-key access control. See the API Reference.
Provider keys: AES-256-GCM at rest (MASTER_KEY, Argon2id-derived). Virtual keys and the admin token: SHA-256 hashed. Outbound SSRF/DNS-rebinding protection. Optional login: WebAuthn passkey, TOTP, and OIDC/GitHub SSO. Details in the Security guide. The repository also ships CrowdSec parsers and scenarios under contrib/crowdsec/ for banning repeated auth failures and rate-limit abuse at the edge; see the CrowdSec wiki page. No prompt or request content is ever logged - see Privacy.
Beyond the shared admin token, provision named dashboard accounts (username + password, optional per-user TOTP) with two roles: admin (full access) or user (scoped by grants: Chat, Usage, Logs, Models, Virtual Keys). Virtual keys belong to a user and per-account rate limits aggregate across their keys. See the Multi-User wiki page.
Prometheus at /metrics (set METRICS_TOKEN so the scrape config carries no admin token). LOG_FORMAT=json emits structured stdout logs for Fluent Bit / Vector / Promtail / Datadog; OTEL_EXPORTER_OTLP_ENDPOINT pushes them to an OTel collector. DEBUG_LOG=true for verbose, DEBUG_LOG_SCOPES=failover,resolve to scope it. See the Configuration wiki.
For pushed alerts rather than scraping, Settings → Alerts POSTs short summaries of operational events (provider down, circuit breaker tripped, failover group out of sync) to a stateless Apprise container that fans them out to Telegram, email, Discord, Slack, Matrix, a webhook, and around 80 other destinations; only the event summary is sent, never request content. See the Alerting wiki.
Backups from the Settings page or POST /api/backups use an unfiltered pg_dump --format=custom with zstd compression (level 12 on request, level 19 for scheduled backups), so the .dump holds every table: providers (encrypted keys), models, virtual key hashes, failover groups and settings, but also request and app logs, the audit log, discovery history, quota snapshots, user accounts, TOTP secrets and recovery-code hashes, and WebAuthn credentials and sessions. Treat a .dump as sensitive.
The dumps are zstd-compressed, so restoring outside the app needs pg_restore 16 or later built with zstd (the postgres:16-alpine image qualifies).
# Direct
pg_restore --clean --if-exists -d YOUR_DB backup_file.dump
# Via Docker
docker exec -i postgres-container pg_restore --clean --if-exists -U user -d dbname < backup_file.dump
Critical requirements: MASTER_KEY must match (provider keys decrypt with it; a mismatch leaves every provider dead). The admin token lives in DATA_DIR/admin-token on the filesystem, not the database (lost → auto-regenerated on next boot). Virtual keys are SHA-256 hashes only - plaintext is never persisted, so lost keys are irrecoverable by design.
MIT. See CONTRIBUTING.md for the contributor license agreement.
Content type
Image
Digest
sha256:c595150e8…
Size
57.8 MB
Last updated
about 15 hours ago
docker pull hugalafutro/model-hotel