VoxCPM2 Wyoming TTS — GPU-accelerated, high-quality diffusion TTS for Home Assistant.
363
Wyoming protocol server wrapping VoxCPM2 (OpenBMB) for Home Assistant voice pipelines. GPU-accelerated via PyTorch/CUDA with torch.compile — no ONNX, no quantization.
VoxCPM2 is a 2B parameter tokenizer-free diffusion autoregressive TTS model trained on 2M+ hours of multilingual speech. It produces high-quality, natural-sounding audio with controllable voice style via inline text descriptions.
docker run --rm --gpus all -p 10300:10300 \
-e VOXCPM_VOICES="Dave:A calm male voice, professional tone|Ella:A warm southern female voice, friendly and clear" \
nullableeth/voxcpm-wyoming:latest
Register in Home Assistant: Settings → Voice Assistants → Wyoming Integration Host: your server IP, Port: whichever host port you mapped.
First startup takes 2–3 minutes — torch.compile JIT-compiles CUDA kernels on the first run and caches them. Subsequent startups are ~30s. The Wyoming port only opens after the warm-up inference completes.
Voices are defined entirely at runtime via VOXCPM_VOICES — nothing is hardcoded in the image.
Format: DisplayName:style description|DisplayName:style description|...
DisplayName — shown in the HA voice selector dropdownstyle description — passed to VoxCPM2 as an inline voice design prefix: (style)textenvironment:
VOXCPM_VOICES: "Dave:A calm male voice, professional tone|Ella:A warm southern female voice, friendly and clear|Nova:An energetic young female voice, upbeat and enthusiastic|Rex:A deep male voice, authoritative and commanding"
If VOXCPM_VOICES is not set, a single Default voice is registered with no style prefix.
Voice consistency: VoxCPM2 is a diffusion model — the same style description produces broadly consistent character but not an identical voice on every run. For consistent voice identity across restarts, use
VOXCPM_REFERENCE_WAV(voice cloning).
Two modes, usable independently or simultaneously:
Reference mode (VoxCPM2 only) — WAV of the target voice, no transcript needed:
environment:
VOXCPM_REFERENCE_WAV: "/voice-refs/my_voice.wav"
volumes:
- /path/to/wavs:/voice-refs:ro
Prompt/continuation mode — WAV + exact transcript for speech context:
environment:
VOXCPM_PROMPT_WAV: "/voice-refs/prompt.wav"
VOXCPM_PROMPT_TEXT: "Exact transcript of the audio in the prompt file."
Both modes are global — applied to every voice in VOXCPM_VOICES. Reference mode anchors timbre; style descriptions apply on top.
| Variable | Default | Description |
|---|---|---|
VOXCPM_VOICES | (empty) | Pipe-delimited Name:style voice definitions |
VOXCPM_MODEL | /models/VoxCPM2 | Model path — baked into image at build time |
VOXCPM_DEVICE | cuda | cuda or cpu |
VOXCPM_OPTIMIZE | true | torch.compile — requires gcc (baked into image). Disable for debugging only |
WYOMING_PORT | 10300 | Internal container port |
OUTPUT_SAMPLE_RATE | 48000 | Output sample rate — 48000 is VoxCPM2 native |
VOXCPM_CFG | 2.0 | Classifier-free guidance strength |
VOXCPM_TIMESTEPS | 10 | Diffusion steps — lower is faster, higher is better quality |
VOXCPM_LOAD_DENOISER | false | Load zipenhancer denoiser (extra VRAM + load time) |
VOXCPM_NORMALIZE | false | Text normalization (numbers, abbreviations) |
VOXCPM_DENOISE_REF | false | Denoise prompt/reference audio — requires VOXCPM_LOAD_DENOISER=true |
VOXCPM_MIN_LEN | 2 | Minimum generation length (tokens) |
VOXCPM_MAX_LEN | 4096 | Maximum generation length (tokens) |
VOXCPM_RETRY_BADCASE | true | Auto-detect and retry low-quality outputs |
VOXCPM_RETRY_MAX_TIMES | 3 | Maximum retry attempts |
VOXCPM_RETRY_RATIO_THRESHOLD | 6.0 | Audio-to-text ratio threshold that triggers retry |
VOXCPM_REFERENCE_WAV | (empty) | Path to reference WAV for voice cloning (no transcript needed) |
VOXCPM_PROMPT_WAV | (empty) | Path to prompt WAV — requires VOXCPM_PROMPT_TEXT |
VOXCPM_PROMPT_TEXT | (empty) | Exact transcript of VOXCPM_PROMPT_WAV |
optimize=true)| Phrase length | Audio duration | Generation time | RTF |
|---|---|---|---|
| Short (8 chars) | 1.76s | 1.40s | 0.80 |
| Medium (35 chars) | 1.76s | 1.28s | 0.73 |
| Long (49 chars) | 3.20s | 2.32s | 0.72 |
RTF < 1.0 = faster than real-time. Minimum latency is ~1.3s regardless of phrase length due to the diffusion architecture.
Timestep tuning:
VOXCPM_TIMESTEPS | Speed | Quality |
|---|---|---|
6 | Fastest — ~0.72 RTF | Good, recommended |
10 | Default — ~0.75 RTF | Better |
25 | Slow — >1.0 RTF | Best |
VRAM usage: ~5.9GB at bfloat16.
First startup: 2–3 minutes (torch.compile kernel compilation). Subsequent startups ~30s from cached kernels.
Comparison with Kokoro: Kokoro (nullableeth/kokoro-wyoming) uses 720MB VRAM and generates short phrases in under 0.5s. VoxCPM2 has a ~1.3s minimum latency floor but produces noticeably better prosody and voice expressiveness. Recommended usage: Kokoro for real-time voice assistant responses, VoxCPM2 for announcements and longer-form content where quality matters more than latency.
wyoming-voxcpm:
container_name: wyoming-voxcpm
image: nullableeth/voxcpm-wyoming:latest
restart: unless-stopped
ports:
- "10211:10300"
environment:
VOXCPM_VOICES: "Dave:A calm male voice, professional tone|Ella:A warm southern female voice, friendly and clear"
VOXCPM_DEVICE: "cuda"
VOXCPM_OPTIMIZE: "true"
VOXCPM_CFG: "2.0"
VOXCPM_TIMESTEPS: "6"
VOXCPM_RETRY_BADCASE: "true"
VOXCPM_REFERENCE_WAV: ""
VOXCPM_PROMPT_WAV: ""
VOXCPM_PROMPT_TEXT: ""
volumes:
- /path/to/voice-refs:/voice-refs:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
run_in_executor) — required for torch.compile CUDA graph consistencytorch.set_float32_matmul_precision('high') enabled for free TF32 tensor core utilizationTORCHINDUCTOR_DISABLE_CUDAGRAPH=1 is not set — CUDA graphs are active and working correctlyhuggingface_hub.snapshot_downloadlocal_files_only=True at runtime — no network access after image build| Image | Purpose | VRAM | Latency |
|---|---|---|---|
nullableeth/whisperx-wyoming | Speech-to-text (WhisperX) | ~1-2GB | — |
nullableeth/kokoro-wyoming | TTS (Kokoro ONNX, fast) | ~720MB | <0.5s |
nullableeth/rvc-wyoming | Voice conversion (RVC) | ~1GB | ~0.5s |
nullableeth/voxcpm-wyoming | TTS (VoxCPM2, high quality) | ~5.9GB | ~1.3s+ |
Content type
Image
Digest
sha256:c12be8f77…
Size
7.2 GB
Last updated
6 months ago
docker pull nullableeth/voxcpm-wyoming