Sign inSign up

sinups/self-hosted-stt-gateway

By sinups

β€’Updated 4 months ago

OpenAI-compatible Whisper STT gateway (turbo/large-v3) with VAD + anti-hallucination filtering

Image
0

1.6K

sinups/self-hosted-stt-gateway repository overview

⁠self-hosted-stt

Self-hosted, OpenAI-compatible speech-to-text. Point any OpenAI client at it and call audio.transcriptions.create(...). Runs locally on your GPU β€” no data leaves your server.

  • 🎯 OpenAI-compatible POST /v1/audio/transcriptions β€” drop-in replacement.
  • ⚑ Whisper turbo (FP16, vLLM) β€” ~real-time, great for in-call transcription.
  • 🎚️ Model switching per request β€” turbo (fast) or large-v3 (accurate).
  • πŸ”‡ No phantom transcripts β€” Silero VAD + loudness gate drop silence, noise and quiet background voices; anti-hallucination filters drop ghost phrases and wrong-language output.
  • πŸ”’ Token auth, language lock (e.g. Russian-only), all via env.
client ──▢ gateway (auth Β· VAD Β· loudness Β· language Β· filters Β· routing)
                 β”œβ”€β”€β–Ά whisper-vllm     (turbo, FP16, GPU)
                 └──▢ faster-whisper   (large-v3, GPU)

⁠Hardware requirements

The turbo model runs on vLLM, which requires an NVIDIA GPU (CUDA). The gateway itself is CPU-only. Plan for the GPU.

MinimumRecommended
GPUNVIDIA, 8 GB VRAM, Turing+ (T4, RTX 20xx)NVIDIA 16–24 GB, Ampere/Ada/Blackwell (A10G, L4, RTX 30/40/50)
Driverβ‰₯ 570 (CUDA 12.8) for the pinned vLLM imagesame
CPU4 cores8 cores
RAM8 GB16 GB
Disk~30 GB (vLLM image ~25 GB + models)40 GB SSD
  • VRAM: FP16 turbo weights are ~1.6 GB; with the 30 s KV-cache and vLLM overhead, 8 GB is comfortable. On a tiny GPU lower VLLM_GPU_MEM (e.g. 0.30) and/or drop the faster-whisper service.
  • Running both turbo + large-v3 on one GPU needs ~10–16 GB.
  • No NVIDIA GPU β†’ no turbo. You can still run large-v3 (faster-whisper) on CPU, but it is not real-time (severalΓ— slower than audio) β€” fine for batch/offline subtitles, not for live calls.
⁠Cloud / where it runs
ProviderInstanceGPUNotes
AWSg4dn.xlargeT4 16 GBcheapest GPU, works (Turing β€” may need VLLM_ATTENTION_BACKEND=TRITON_ATTN)
AWSg5.xlargeA10G 24 GBrecommended, fast
AWSg6.xlargeL4 24 GBnewest, efficient
GCPn1 + T4 / L4, g2T4 / L4works
AzureNCas_T4_v3T4 16 GBworks
OthersRunPod / Vast.ai / Lambda / Hetzner GPUany NVIDIAworks, cheapest for testing

❌ Shared hosting / a regular $5 VPS / AWS t3/m5 (no GPU) will NOT run turbo. ❌ AWS g4ad (AMD GPU) is not supported β€” must be NVIDIA. For a live-call use case you need one of the GPU instances above.

⁠Software

Verify the GPU is visible to Docker:

docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

⁠Quickstart (3 commands)

git clone https://github.com/sinups/self-hosted-stt.git
cd self-hosted-stt
cp .env.example .env          # then edit STT_TOKEN (openssl rand -hex 24)
docker compose up -d

First start downloads the model weights (~1.6 GB turbo, ~3 GB large-v3) β€” give it a couple of minutes. Watch progress:

docker compose logs -f gateway whisper-vllm
⁠Faster path β€” pre-built gateway image

Skip building the gateway by using the published image from Docker Hub (sinups/self-hosted-stt-gateway⁠):

cp .env.example .env          # set STT_TOKEN
docker compose -f docker-compose.images.yml up -d

(The small whisper-vllm layer is still built locally β€” it's quick.)

Test it:

curl -s http://localhost:8080/v1/audio/transcriptions \
  -H "Authorization: Bearer $STT_TOKEN" \
  -F [email protected] -F model=turbo
# -> {"text":"..."}   (or {"text":""} for silence/noise/quiet background)

⁠Use it (OpenAI clients)

Python

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="YOUR_STT_TOKEN")
r = client.audio.transcriptions.create(
        model="turbo", file=open("sample.wav", "rb"))
print(r.text)

Node

import OpenAI from "openai";
import fs from "fs";
const client = new OpenAI({ baseURL: "http://localhost:8080/v1", apiKey: "YOUR_STT_TOKEN" });
const r = await client.audio.transcriptions.create({
  model: "turbo", file: fs.createReadStream("sample.wav") });
console.log(r.text);
⁠Request fields
FieldNotes
fileaudio (WAV 16 kHz mono ideal; mp3/webm/ogg/m4a also accepted)
modelturbo (fast) or large-v3 (accurate). Empty β†’ DEFAULT_MODEL
languagee.g. ru. Ignored if STT_FORCE_LANGUAGE is set
response_formatjson (default) or text

Response headers expose gate diagnostics: x-stt-gate (ok/no_speech/too_quiet), x-stt-dbfs, x-stt-speech-ms, x-stt-provider.

An empty {"text":""} is normal β€” it means the gate filtered silence, noise or quiet background. Don't fall back to another STT on empty: that re-introduces the phantom transcripts this service exists to prevent.


⁠Configuration (.env)

VariableDefaultDescription
STT_TOKENβ€”Bearer token(s), comma-separated. Empty = no auth
GATEWAY_PORT8080Host port for the API
DEFAULT_MODELturboModel when client omits it
STT_FORCE_LANGUAGE(off)Lock to one language, e.g. ru. Ignores request language incl. auto
STT_DEFAULT_LANGUAGEruUsed when request omits language and FORCE is off ("" = auto)
ENABLE_GATEtrueVAD + loudness gate
VAD_THRESHOLD0.6Higher = stricter speech detection
MIN_SPEECH_MS250Min speech per chunk; lower (~150) to catch short "yes/no"
MIN_SPEECH_DBFS-36Speech quieter than this (background) β†’ dropped. Lower (-40) if your speaker is quiet; raise (-32) if background leaks
VLLM_GPU_MEM0.45GPU fraction for vLLM (lower if OOM)

After editing .env: docker compose up -d.

⁠Russian-only example
STT_FORCE_LANGUAGE=ru

Guarantees Russian output β€” no English/foreign phantom transcripts.


⁠Tuning for in-call transcription

  • Send 5–15 s chunks, ideally with 1–2 s overlap or cut on speech pauses (hard fixed cuts split words/names across chunks).
  • Prefer one audio stream per participant (Whisper has no diarization).
  • Calibrate MIN_SPEECH_DBFS from the x-stt-dbfs header on real call audio: set it between your speaker's level and the background level.

⁠Troubleshooting

SymptomFix
unsatisfied condition: cuda>=13.0 on startYour driver is < 580. The image is pinned to vLLM v0.10.2 (CUDA 12.8) for exactly this β€” make sure you didn't bump it. Driver < 570? Use an older vLLM tag.
Int8 not supported for this architectureINT8 quantized Whisper isn't supported on Blackwell (RTX 50xx) with CUDA 12.8 β€” this repo uses FP16 turbo (openai/whisper-large-v3-turbo) on purpose.
Please install vllm[audio]The whisper-vllm image bakes librosa/soundfile; rebuild it (docker compose build whisper-vllm).
OOM on a small GPULower VLLM_GPU_MEM (e.g. 0.30) and/or drop the faster-whisper service.
Only want turboDelete the faster-whisper service block in docker-compose.yml.

⁠CI/CD (GitHub Actions β†’ Docker Hub)

.github/workflows/docker.yml builds and pushes the gateway image to Docker Hub automatically:

TriggerImage tags pushed
push to masterlatest, sha-<short>
push tag vX.Y.ZX.Y.Z, X.Y, latest
manual (Actions β†’ Run workflow)as above

One-time setup β€” repo secrets (Settings β†’ Secrets and variables β†’ Actions):

SecretValue
DOCKERHUB_USERNAMEyour Docker Hub user (e.g. sinups)
DOCKERHUB_TOKENa Docker Hub access token (Account β†’ Security β†’ New Access Token)

Or via CLI:

gh secret set DOCKERHUB_USERNAME --body "sinups"
gh secret set DOCKERHUB_TOKEN    --body "dckr_pat_xxx"
⁠Releasing a version
git tag v0.2.0
git push origin v0.2.0      # CI builds & pushes sinups/self-hosted-stt-gateway:0.2.0

Then bump the tag in docker-compose.images.yml (or keep :latest).

⁠License

MIT

Tag summary

Content type

Image

Digest

sha256:34a897565…

Size

594.6 MB

Last updated

4 months ago

docker pull sinups/self-hosted-stt-gateway