High-accuracy speech-to-text with WhisperX. 3-4x faster than faster-whisper. Wyoming for HA.
1.7K
Wyoming protocol server for WhisperX with superior accuracy and speed. Integrates seamlessly with Home Assistant for voice assistants.
Speech-to-text server using WhisperX with word-level timestamps, voice activity detection, and improved accuracy over standard Whisper implementations. Up to 3-4x faster than faster-whisper after warm-up.
GPU mode (recommended):
docker run -d \
--name whisperx \
--gpus all \
-p 10300:10300 \
--restart unless-stopped \
nullableeth/whisperx-wyoming:latest
CPU mode:
docker run -d \
--name whisperx \
-p 10300:10300 \
-e WHISPER_DEVICE=cpu \
-e WHISPER_COMPUTE_TYPE=int8 \
--restart unless-stopped \
nullableeth/whisperx-wyoming:latest
Configure via environment variables - no need to override the command:
| Variable | Default | Options | Description |
|---|---|---|---|
WYOMING_URI | tcp://0.0.0.0:10300 | tcp://host:port | Wyoming server bind address |
WHISPER_MODEL | base | tiny, base, small, medium, large-v2, large-v3 | Model size (accuracy vs speed) |
WHISPER_LANGUAGE | en | See languages | Primary language code |
WHISPER_DEVICE | cuda | cuda, cpu | Device to run inference on |
WHISPER_COMPUTE_TYPE | float16 | float16, int8 | Precision (GPU: float16, CPU: int8) |
services:
whisperx:
image: nullableeth/whisperx-wyoming:latest
container_name: whisperx
restart: unless-stopped
ports:
- "10300:10300"
environment:
WHISPER_MODEL: "base" # tiny, base, small, medium, large-v2
WHISPER_LANGUAGE: "en" # Language code
WHISPER_DEVICE: "cuda" # cuda or cpu
WHISPER_COMPUTE_TYPE: "float16" # float16 or int8
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
You can also use command-line arguments (env vars take precedence if both are set):
docker run -d \
--name whisperx \
--gpus all \
-p 10300:10300 \
nullableeth/whisperx-wyoming:latest \
--model medium \
--language en \
--device cuda \
--compute-type float16
<docker-host-ip>, Port: 10300Choose based on your accuracy vs speed needs:
| Model | Size | VRAM | Accuracy | Speed | Best For |
|---|---|---|---|---|---|
tiny | 39M | ~1GB | ⭐⭐ | ⚡⚡⚡⚡⚡ | Testing, low-resource |
base | 74M | ~1GB | ⭐⭐⭐ | ⚡⚡⚡⚡ | Default, balanced |
small | 244M | ~2GB | ⭐⭐⭐⭐ | ⚡⚡⚡ | Better accuracy |
medium | 769M | ~5GB | ⭐⭐⭐⭐⭐ | ⚡⚡ | High accuracy |
large-v2 | 1550M | ~10GB | ⭐⭐⭐⭐⭐⭐ | ⚡ | Best accuracy |
large-v3 | 1550M | ~10GB | ⭐⭐⭐⭐⭐⭐ | ⚡ | Latest, best overall |
Recommendation: Start with base, upgrade to small or medium if accuracy isn't sufficient.
Full language support via ISO 639-1 codes:
| Code | Language | Code | Language | Code | Language |
|---|---|---|---|---|---|
en | English | es | Spanish | fr | French |
de | German | it | Italian | pt | Portuguese |
nl | Dutch | pl | Polish | tr | Turkish |
ru | Russian | cs | Czech | ar | Arabic |
zh | Chinese | ja | Japanese | ko | Korean |
hi | Hindi | hu | Hungarian | sv | Swedish |
da | Danish | no | Norwegian | fi | Finnish |
50+ languages total. Full list: https://github.com/openai/whisper#available-models-and-languages
Set via WHISPER_LANGUAGE environment variable or --language flag.
| Type | Description | Use Case |
|---|---|---|
float16 | Half-precision floating point | GPU inference (default) |
int8 | 8-bit integer quantization | CPU inference, lower VRAM |
float32 | Full precision | Not recommended (slower, more VRAM) |
Recommendation: Use float16 for GPU, int8 for CPU.
GPU Mode (RTX 4070 + base model):
GPU Mode (RTX 4070 + medium model):
CPU Mode (Intel i7):
Comparison to faster-whisper:
environment:
WHISPER_MODEL: "large-v3"
WHISPER_LANGUAGE: "en"
WHISPER_DEVICE: "cuda"
WHISPER_COMPUTE_TYPE: "float16"
Requires ~10GB VRAM
environment:
WHISPER_MODEL: "base"
WHISPER_LANGUAGE: "en"
WHISPER_DEVICE: "cuda"
WHISPER_COMPUTE_TYPE: "float16"
Requires ~2GB VRAM
environment:
WHISPER_MODEL: "base"
WHISPER_LANGUAGE: "en"
WHISPER_DEVICE: "cpu"
WHISPER_COMPUTE_TYPE: "int8"
# Remove deploy.resources section
No GPU required
environment:
WHISPER_MODEL: "small"
WHISPER_LANGUAGE: "es" # Spanish
WHISPER_DEVICE: "cuda"
WHISPER_COMPUTE_TYPE: "float16"
Auto-detects if language doesn't match
GPU (Recommended):
--gpus all)CPU (Fallback):
Clean, timestamped logs for easy debugging:
2026-03-20 12:00:00 - INFO - Loading WhisperX base on cuda...
2026-03-20 12:00:02 - INFO - ✓ Model loaded
2026-03-20 12:00:02 - INFO - Starting server on tcp://0.0.0.0:10300
2026-03-20 12:00:15 - INFO - Transcribing 3.98s of audio...
2026-03-20 12:00:15 - INFO - ✓ Transcribed in 0.44s: 'turn on the living room lights'
Normal behavior! VAD model loads on first use (~2s overhead). All subsequent transcriptions are fast.
# Verify NVIDIA Docker support
docker run --rm --gpus all nvidia/cuda:12.1.0-base nvidia-smi
# Check container has GPU access
docker exec whisperx nvidia-smi
# Verify CUDA in container
docker exec whisperx python3 -c "import torch; print('CUDA:', torch.cuda.is_available())"
base → small → mediumWHISPER_LANGUAGE to match spoken languagemedium → small → baseWHISPER_COMPUTE_TYPE=int8WHISPER_DEVICE=cpu (slower but no VRAM)# Check logs
docker logs whisperx
# Common issues:
# 1. Port already in use → change port
# 2. No GPU access → add --gpus all
# 3. CUDA version mismatch → update nvidia-docker
vs faster-whisper:
vs OpenAI Whisper:
vs Vosk:
From community testing on LibriSpeech dataset:
| Model | WER (%) | CER (%) | Latency (ms) |
|---|---|---|---|
| WhisperX (fp16, B=16) | 9.99 | 3.6 | 1,170-2,660 |
| FasterWhisper (fp16) | 11.98 | 4.7 | 2,580 |
| OpenAI Whisper | 10.8 | 4.3 | 10,800 |
Lower is better for WER/CER, latency varies by audio length
git clone <your-repo>
cd whisperx-wyoming
docker build -t whisperx-wyoming:latest .
For issues specific to this container, please open an issue on GitHub. For WhisperX questions, see the WhisperX documentation.
Content type
Image
Digest
sha256:a8c7004d3…
Size
6.3 GB
Last updated
7 months ago
docker pull nullableeth/whisperx-wyoming