Sign inSign up

nullableeth/whisperx-wyoming

By nullableeth

•Updated 7 months ago

High-accuracy speech-to-text with WhisperX. 3-4x faster than faster-whisper. Wyoming for HA.

Image
Machine learning & AI
0

1.7K

nullableeth/whisperx-wyoming repository overview

⁠WhisperX Wyoming - High-Accuracy Speech-to-Text

Wyoming protocol server for WhisperX with superior accuracy and speed. Integrates seamlessly with Home Assistant for voice assistants.

⁠What This Is

Speech-to-text server using WhisperX with word-level timestamps, voice activity detection, and improved accuracy over standard Whisper implementations. Up to 3-4x faster than faster-whisper after warm-up.

⁠Features

  • ✅ Superior accuracy - 9.99% WER vs 11.98% for faster-whisper
  • 🚀 Fast inference - 0.4-0.6s transcription after warm-up
  • 🎯 Word-level timestamps - Precise alignment
  • 🔊 VAD built-in - Voice activity detection filters noise
  • 🌍 Multi-language - Supports all Whisper languages
  • 🏠 Home Assistant ready - Wyoming protocol integration

⁠Quick Start

GPU mode (recommended):

docker run -d \
  --name whisperx \
  --gpus all \
  -p 10300:10300 \
  --restart unless-stopped \
  nullableeth/whisperx-wyoming:latest

CPU mode:

docker run -d \
  --name whisperx \
  -p 10300:10300 \
  -e WHISPER_DEVICE=cpu \
  -e WHISPER_COMPUTE_TYPE=int8 \
  --restart unless-stopped \
  nullableeth/whisperx-wyoming:latest

⁠Configuration

Configure via environment variables - no need to override the command:

VariableDefaultOptionsDescription
WYOMING_URItcp://0.0.0.0:10300tcp://host:portWyoming server bind address
WHISPER_MODELbasetiny, base, small, medium, large-v2, large-v3Model size (accuracy vs speed)
WHISPER_LANGUAGEenSee languages⁠Primary language code
WHISPER_DEVICEcudacuda, cpuDevice to run inference on
WHISPER_COMPUTE_TYPEfloat16float16, int8Precision (GPU: float16, CPU: int8)
services:
  whisperx:
    image: nullableeth/whisperx-wyoming:latest
    container_name: whisperx
    restart: unless-stopped
    ports:
      - "10300:10300"
    environment:
      WHISPER_MODEL: "base"          # tiny, base, small, medium, large-v2
      WHISPER_LANGUAGE: "en"         # Language code
      WHISPER_DEVICE: "cuda"         # cuda or cpu
      WHISPER_COMPUTE_TYPE: "float16" # float16 or int8
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
⁠Command-Line Arguments (Alternative)

You can also use command-line arguments (env vars take precedence if both are set):

docker run -d \
  --name whisperx \
  --gpus all \
  -p 10300:10300 \
  nullableeth/whisperx-wyoming:latest \
  --model medium \
  --language en \
  --device cuda \
  --compute-type float16

⁠Home Assistant Integration

  1. Settings → Devices & Services → Add Integration
  2. Search "Wyoming Protocol"
  3. Host: <docker-host-ip>, Port: 10300
  4. Configure voice assistant to use WhisperX for speech-to-text

⁠Available Models

Choose based on your accuracy vs speed needs:

ModelSizeVRAMAccuracySpeedBest For
tiny39M~1GB⭐⭐⚡⚡⚡⚡⚡Testing, low-resource
base74M~1GB⭐⭐⭐⚡⚡⚡⚡Default, balanced
small244M~2GB⭐⭐⭐⭐⚡⚡⚡Better accuracy
medium769M~5GB⭐⭐⭐⭐⭐⚡⚡High accuracy
large-v21550M~10GB⭐⭐⭐⭐⭐⭐⚡Best accuracy
large-v31550M~10GB⭐⭐⭐⭐⭐⭐⚡Latest, best overall

Recommendation: Start with base, upgrade to small or medium if accuracy isn't sufficient.

⁠Supported Languages

Full language support via ISO 639-1 codes:

CodeLanguageCodeLanguageCodeLanguage
enEnglishesSpanishfrFrench
deGermanitItalianptPortuguese
nlDutchplPolishtrTurkish
ruRussiancsCzecharArabic
zhChinesejaJapanesekoKorean
hiHindihuHungariansvSwedish
daDanishnoNorwegianfiFinnish

50+ languages total. Full list: https://github.com/openai/whisper#available-models-and-languages⁠

Set via WHISPER_LANGUAGE environment variable or --language flag.

⁠Compute Types

TypeDescriptionUse Case
float16Half-precision floating pointGPU inference (default)
int88-bit integer quantizationCPU inference, lower VRAM
float32Full precisionNot recommended (slower, more VRAM)

Recommendation: Use float16 for GPU, int8 for CPU.

⁠Performance

GPU Mode (RTX 4070 + base model):

  • First transcription: ~2.5s (includes VAD model loading)
  • Subsequent: ~0.4-0.6s per utterance
  • VRAM usage: ~2-3GB

GPU Mode (RTX 4070 + medium model):

  • First transcription: ~3s
  • Subsequent: ~0.8-1.2s per utterance
  • VRAM usage: ~5GB

CPU Mode (Intel i7):

  • First transcription: ~8-12s
  • Subsequent: ~3-5s per utterance
  • RAM usage: ~4GB

Comparison to faster-whisper:

  • ✅ 3-4x faster on subsequent transcriptions
  • ✅ Better accuracy - correctly handles technical terms, proper nouns
  • ⚠️ 2s slower on first transcription (VAD initialization)

⁠Example Configurations

⁠Maximum Accuracy (Large GPU)
environment:
  WHISPER_MODEL: "large-v3"
  WHISPER_LANGUAGE: "en"
  WHISPER_DEVICE: "cuda"
  WHISPER_COMPUTE_TYPE: "float16"

Requires ~10GB VRAM

environment:
  WHISPER_MODEL: "base"
  WHISPER_LANGUAGE: "en"
  WHISPER_DEVICE: "cuda"
  WHISPER_COMPUTE_TYPE: "float16"

Requires ~2GB VRAM

⁠CPU-Only Mode
environment:
  WHISPER_MODEL: "base"
  WHISPER_LANGUAGE: "en"
  WHISPER_DEVICE: "cpu"
  WHISPER_COMPUTE_TYPE: "int8"
# Remove deploy.resources section

No GPU required

⁠Multi-Language
environment:
  WHISPER_MODEL: "small"
  WHISPER_LANGUAGE: "es"  # Spanish
  WHISPER_DEVICE: "cuda"
  WHISPER_COMPUTE_TYPE: "float16"

Auto-detects if language doesn't match

⁠Requirements

GPU (Recommended):

  • NVIDIA GPU with CUDA 12.1+ support
  • Minimum 2GB VRAM (base model)
  • 5GB+ VRAM for medium/large models
  • Docker with NVIDIA Container Toolkit (--gpus all)

CPU (Fallback):

  • 4GB+ RAM
  • Multi-core CPU recommended
  • 5-10x slower than GPU

⁠Logs

Clean, timestamped logs for easy debugging:

2026-03-20 12:00:00 - INFO - Loading WhisperX base on cuda...
2026-03-20 12:00:02 - INFO - ✓ Model loaded
2026-03-20 12:00:02 - INFO - Starting server on tcp://0.0.0.0:10300
2026-03-20 12:00:15 - INFO - Transcribing 3.98s of audio...
2026-03-20 12:00:15 - INFO - ✓ Transcribed in 0.44s: 'turn on the living room lights'

⁠Troubleshooting

⁠Slow first transcription

Normal behavior! VAD model loads on first use (~2s overhead). All subsequent transcriptions are fast.

⁠GPU not detected
# Verify NVIDIA Docker support
docker run --rm --gpus all nvidia/cuda:12.1.0-base nvidia-smi

# Check container has GPU access
docker exec whisperx nvidia-smi

# Verify CUDA in container
docker exec whisperx python3 -c "import torch; print('CUDA:', torch.cuda.is_available())"
⁠Poor accuracy
  • Try larger model: Upgrade from base → small → medium
  • Check language: Set WHISPER_LANGUAGE to match spoken language
  • Audio quality: Ensure clean audio input (16kHz recommended)
  • Microphone levels: Verify levels in Home Assistant
⁠Empty transcriptions
  • Audio too short: Minimum ~0.5s needed
  • Microphone muted: Check HA microphone settings
  • Wyoming config: Verify integration points to correct port
⁠High VRAM usage
  • Use smaller model: Switch from medium → small → base
  • Use int8: Set WHISPER_COMPUTE_TYPE=int8
  • CPU mode: Set WHISPER_DEVICE=cpu (slower but no VRAM)
⁠Container won't start
# Check logs
docker logs whisperx

# Common issues:
# 1. Port already in use → change port
# 2. No GPU access → add --gpus all
# 3. CUDA version mismatch → update nvidia-docker

⁠Comparison to Alternatives

vs faster-whisper:

  • ✅ WhisperX: 3-4x faster after warm-up
  • ✅ WhisperX: Better accuracy (9.99% vs 11.98% WER)
  • ⚠️ faster-whisper: Faster cold start (~500ms less)

vs OpenAI Whisper:

  • ✅ WhisperX: Much faster (batched processing)
  • ✅ WhisperX: Word-level timestamps
  • ✅ WhisperX: Built-in VAD

vs Vosk:

  • ✅ WhisperX: Better accuracy
  • ✅ Vosk: Lower resource usage (~500MB RAM)
  • ✅ WhisperX: Better with accents/dialects

⁠Known Issues

  • First transcription takes ~2s extra for VAD initialization (normal)
  • Lightning checkpoint upgrade message on startup (harmless, can be ignored)
  • Requires GPU for optimal performance (CPU works but 5-10x slower)

⁠Benchmark Results

From community testing on LibriSpeech dataset:

ModelWER (%)CER (%)Latency (ms)
WhisperX (fp16, B=16)9.993.61,170-2,660
FasterWhisper (fp16)11.984.72,580
OpenAI Whisper10.84.310,800

Lower is better for WER/CER, latency varies by audio length

⁠Building From Source

git clone <your-repo>
cd whisperx-wyoming
docker build -t whisperx-wyoming:latest .

⁠License

  • WhisperX: MIT License
  • Whisper models: Apache 2.0
  • Wyoming Protocol: MIT License

⁠Support

For issues specific to this container, please open an issue on GitHub. For WhisperX questions, see the WhisperX documentation⁠.

Tag summary

Content type

Image

Digest

sha256:a8c7004d3…

Size

6.3 GB

Last updated

7 months ago

docker pull nullableeth/whisperx-wyoming