OpenAI-compatible Whisper STT gateway (turbo/large-v3) with VAD + anti-hallucination filtering
1.6K
Self-hosted, OpenAI-compatible speech-to-text. Point any OpenAI client at it
and call audio.transcriptions.create(...). Runs locally on your GPU β no data
leaves your server.
POST /v1/audio/transcriptions β drop-in replacement.turbo (fast) or large-v3 (accurate).client βββΆ gateway (auth Β· VAD Β· loudness Β· language Β· filters Β· routing)
ββββΆ whisper-vllm (turbo, FP16, GPU)
ββββΆ faster-whisper (large-v3, GPU)
The turbo model runs on vLLM, which requires an NVIDIA GPU (CUDA). The
gateway itself is CPU-only. Plan for the GPU.
| Minimum | Recommended | |
|---|---|---|
| GPU | NVIDIA, 8 GB VRAM, Turing+ (T4, RTX 20xx) | NVIDIA 16β24 GB, Ampere/Ada/Blackwell (A10G, L4, RTX 30/40/50) |
| Driver | β₯ 570 (CUDA 12.8) for the pinned vLLM image | same |
| CPU | 4 cores | 8 cores |
| RAM | 8 GB | 16 GB |
| Disk | ~30 GB (vLLM image ~25 GB + models) | 40 GB SSD |
VLLM_GPU_MEM (e.g.
0.30) and/or drop the faster-whisper service.turbo + large-v3 on one GPU needs ~10β16 GB.turbo. You can still run large-v3 (faster-whisper)
on CPU, but it is not real-time (severalΓ slower than audio) β fine for
batch/offline subtitles, not for live calls.| Provider | Instance | GPU | Notes |
|---|---|---|---|
| AWS | g4dn.xlarge | T4 16 GB | cheapest GPU, works (Turing β may need VLLM_ATTENTION_BACKEND=TRITON_ATTN) |
| AWS | g5.xlarge | A10G 24 GB | recommended, fast |
| AWS | g6.xlarge | L4 24 GB | newest, efficient |
| GCP | n1 + T4 / L4, g2 | T4 / L4 | works |
| Azure | NCas_T4_v3 | T4 16 GB | works |
| Others | RunPod / Vast.ai / Lambda / Hetzner GPU | any NVIDIA | works, cheapest for testing |
β Shared hosting / a regular $5 VPS / AWS
t3/m5(no GPU) will NOT runturbo. β AWSg4ad(AMD GPU) is not supported β must be NVIDIA. For a live-call use case you need one of the GPU instances above.
nvidia-smi (top-right). Driver < 570 or
β₯ 580: see Troubleshootingβ .Verify the GPU is visible to Docker:
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
git clone https://github.com/sinups/self-hosted-stt.git
cd self-hosted-stt
cp .env.example .env # then edit STT_TOKEN (openssl rand -hex 24)
docker compose up -d
First start downloads the model weights (~1.6 GB turbo, ~3 GB large-v3) β give it a couple of minutes. Watch progress:
docker compose logs -f gateway whisper-vllm
Skip building the gateway by using the published image from Docker Hub
(sinups/self-hosted-stt-gatewayβ ):
cp .env.example .env # set STT_TOKEN
docker compose -f docker-compose.images.yml up -d
(The small whisper-vllm layer is still built locally β it's quick.)
Test it:
curl -s http://localhost:8080/v1/audio/transcriptions \
-H "Authorization: Bearer $STT_TOKEN" \
-F [email protected] -F model=turbo
# -> {"text":"..."} (or {"text":""} for silence/noise/quiet background)
Python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="YOUR_STT_TOKEN")
r = client.audio.transcriptions.create(
model="turbo", file=open("sample.wav", "rb"))
print(r.text)
Node
import OpenAI from "openai";
import fs from "fs";
const client = new OpenAI({ baseURL: "http://localhost:8080/v1", apiKey: "YOUR_STT_TOKEN" });
const r = await client.audio.transcriptions.create({
model: "turbo", file: fs.createReadStream("sample.wav") });
console.log(r.text);
| Field | Notes |
|---|---|
file | audio (WAV 16 kHz mono ideal; mp3/webm/ogg/m4a also accepted) |
model | turbo (fast) or large-v3 (accurate). Empty β DEFAULT_MODEL |
language | e.g. ru. Ignored if STT_FORCE_LANGUAGE is set |
response_format | json (default) or text |
Response headers expose gate diagnostics: x-stt-gate (ok/no_speech/too_quiet),
x-stt-dbfs, x-stt-speech-ms, x-stt-provider.
An empty
{"text":""}is normal β it means the gate filtered silence, noise or quiet background. Don't fall back to another STT on empty: that re-introduces the phantom transcripts this service exists to prevent.
.env)| Variable | Default | Description |
|---|---|---|
STT_TOKEN | β | Bearer token(s), comma-separated. Empty = no auth |
GATEWAY_PORT | 8080 | Host port for the API |
DEFAULT_MODEL | turbo | Model when client omits it |
STT_FORCE_LANGUAGE | (off) | Lock to one language, e.g. ru. Ignores request language incl. auto |
STT_DEFAULT_LANGUAGE | ru | Used when request omits language and FORCE is off ("" = auto) |
ENABLE_GATE | true | VAD + loudness gate |
VAD_THRESHOLD | 0.6 | Higher = stricter speech detection |
MIN_SPEECH_MS | 250 | Min speech per chunk; lower (~150) to catch short "yes/no" |
MIN_SPEECH_DBFS | -36 | Speech quieter than this (background) β dropped. Lower (-40) if your speaker is quiet; raise (-32) if background leaks |
VLLM_GPU_MEM | 0.45 | GPU fraction for vLLM (lower if OOM) |
After editing .env: docker compose up -d.
STT_FORCE_LANGUAGE=ru
Guarantees Russian output β no English/foreign phantom transcripts.
MIN_SPEECH_DBFS from the x-stt-dbfs header on real call audio:
set it between your speaker's level and the background level.| Symptom | Fix |
|---|---|
unsatisfied condition: cuda>=13.0 on start | Your driver is < 580. The image is pinned to vLLM v0.10.2 (CUDA 12.8) for exactly this β make sure you didn't bump it. Driver < 570? Use an older vLLM tag. |
Int8 not supported for this architecture | INT8 quantized Whisper isn't supported on Blackwell (RTX 50xx) with CUDA 12.8 β this repo uses FP16 turbo (openai/whisper-large-v3-turbo) on purpose. |
Please install vllm[audio] | The whisper-vllm image bakes librosa/soundfile; rebuild it (docker compose build whisper-vllm). |
| OOM on a small GPU | Lower VLLM_GPU_MEM (e.g. 0.30) and/or drop the faster-whisper service. |
| Only want turbo | Delete the faster-whisper service block in docker-compose.yml. |
.github/workflows/docker.yml builds and pushes the gateway image to
Docker Hub automatically:
| Trigger | Image tags pushed |
|---|---|
push to master | latest, sha-<short> |
push tag vX.Y.Z | X.Y.Z, X.Y, latest |
| manual (Actions β Run workflow) | as above |
One-time setup β repo secrets (Settings β Secrets and variables β Actions):
| Secret | Value |
|---|---|
DOCKERHUB_USERNAME | your Docker Hub user (e.g. sinups) |
DOCKERHUB_TOKEN | a Docker Hub access token (Account β Security β New Access Token) |
Or via CLI:
gh secret set DOCKERHUB_USERNAME --body "sinups"
gh secret set DOCKERHUB_TOKEN --body "dckr_pat_xxx"
git tag v0.2.0
git push origin v0.2.0 # CI builds & pushes sinups/self-hosted-stt-gateway:0.2.0
Then bump the tag in docker-compose.images.yml (or keep :latest).
MIT
Content type
Image
Digest
sha256:34a897565β¦
Size
594.6 MB
Last updated
4 months ago
docker pull sinups/self-hosted-stt-gateway