Sign inSign up

nullableeth/voxcpm-wyoming

By nullableeth

•Updated 6 months ago

VoxCPM2 Wyoming TTS — GPU-accelerated, high-quality diffusion TTS for Home Assistant.

Image
Machine learning & AI
0

363

nullableeth/voxcpm-wyoming repository overview

⁠nullableeth/voxcpm-wyoming

Wyoming protocol server wrapping VoxCPM2⁠ (OpenBMB) for Home Assistant voice pipelines. GPU-accelerated via PyTorch/CUDA with torch.compile — no ONNX, no quantization.

VoxCPM2 is a 2B parameter tokenizer-free diffusion autoregressive TTS model trained on 2M+ hours of multilingual speech. It produces high-quality, natural-sounding audio with controllable voice style via inline text descriptions.


⁠Quick start

docker run --rm --gpus all -p 10300:10300 \
  -e VOXCPM_VOICES="Dave:A calm male voice, professional tone|Ella:A warm southern female voice, friendly and clear" \
  nullableeth/voxcpm-wyoming:latest

Register in Home Assistant: Settings → Voice Assistants → Wyoming Integration Host: your server IP, Port: whichever host port you mapped.

First startup takes 2–3 minutes — torch.compile JIT-compiles CUDA kernels on the first run and caches them. Subsequent startups are ~30s. The Wyoming port only opens after the warm-up inference completes.


⁠Named voices

Voices are defined entirely at runtime via VOXCPM_VOICES — nothing is hardcoded in the image.

Format: DisplayName:style description|DisplayName:style description|...

  • DisplayName — shown in the HA voice selector dropdown
  • style description — passed to VoxCPM2 as an inline voice design prefix: (style)text
environment:
  VOXCPM_VOICES: "Dave:A calm male voice, professional tone|Ella:A warm southern female voice, friendly and clear|Nova:An energetic young female voice, upbeat and enthusiastic|Rex:A deep male voice, authoritative and commanding"

If VOXCPM_VOICES is not set, a single Default voice is registered with no style prefix.

Voice consistency: VoxCPM2 is a diffusion model — the same style description produces broadly consistent character but not an identical voice on every run. For consistent voice identity across restarts, use VOXCPM_REFERENCE_WAV (voice cloning).


⁠Voice cloning

Two modes, usable independently or simultaneously:

Reference mode (VoxCPM2 only) — WAV of the target voice, no transcript needed:

environment:
  VOXCPM_REFERENCE_WAV: "/voice-refs/my_voice.wav"
volumes:
  - /path/to/wavs:/voice-refs:ro

Prompt/continuation mode — WAV + exact transcript for speech context:

environment:
  VOXCPM_PROMPT_WAV: "/voice-refs/prompt.wav"
  VOXCPM_PROMPT_TEXT: "Exact transcript of the audio in the prompt file."

Both modes are global — applied to every voice in VOXCPM_VOICES. Reference mode anchors timbre; style descriptions apply on top.


⁠All environment variables

VariableDefaultDescription
VOXCPM_VOICES(empty)Pipe-delimited Name:style voice definitions
VOXCPM_MODEL/models/VoxCPM2Model path — baked into image at build time
VOXCPM_DEVICEcudacuda or cpu
VOXCPM_OPTIMIZEtruetorch.compile — requires gcc (baked into image). Disable for debugging only
WYOMING_PORT10300Internal container port
OUTPUT_SAMPLE_RATE48000Output sample rate — 48000 is VoxCPM2 native
VOXCPM_CFG2.0Classifier-free guidance strength
VOXCPM_TIMESTEPS10Diffusion steps — lower is faster, higher is better quality
VOXCPM_LOAD_DENOISERfalseLoad zipenhancer denoiser (extra VRAM + load time)
VOXCPM_NORMALIZEfalseText normalization (numbers, abbreviations)
VOXCPM_DENOISE_REFfalseDenoise prompt/reference audio — requires VOXCPM_LOAD_DENOISER=true
VOXCPM_MIN_LEN2Minimum generation length (tokens)
VOXCPM_MAX_LEN4096Maximum generation length (tokens)
VOXCPM_RETRY_BADCASEtrueAuto-detect and retry low-quality outputs
VOXCPM_RETRY_MAX_TIMES3Maximum retry attempts
VOXCPM_RETRY_RATIO_THRESHOLD6.0Audio-to-text ratio threshold that triggers retry
VOXCPM_REFERENCE_WAV(empty)Path to reference WAV for voice cloning (no transcript needed)
VOXCPM_PROMPT_WAV(empty)Path to prompt WAV — requires VOXCPM_PROMPT_TEXT
VOXCPM_PROMPT_TEXT(empty)Exact transcript of VOXCPM_PROMPT_WAV

⁠Performance (RTX 4070 Super, optimize=true)

Phrase lengthAudio durationGeneration timeRTF
Short (8 chars)1.76s1.40s0.80
Medium (35 chars)1.76s1.28s0.73
Long (49 chars)3.20s2.32s0.72

RTF < 1.0 = faster than real-time. Minimum latency is ~1.3s regardless of phrase length due to the diffusion architecture.

Timestep tuning:

VOXCPM_TIMESTEPSSpeedQuality
6Fastest — ~0.72 RTFGood, recommended
10Default — ~0.75 RTFBetter
25Slow — >1.0 RTFBest

VRAM usage: ~5.9GB at bfloat16.

First startup: 2–3 minutes (torch.compile kernel compilation). Subsequent startups ~30s from cached kernels.

Comparison with Kokoro: Kokoro (nullableeth/kokoro-wyoming) uses 720MB VRAM and generates short phrases in under 0.5s. VoxCPM2 has a ~1.3s minimum latency floor but produces noticeably better prosody and voice expressiveness. Recommended usage: Kokoro for real-time voice assistant responses, VoxCPM2 for announcements and longer-form content where quality matters more than latency.


⁠docker-compose example

wyoming-voxcpm:
  container_name: wyoming-voxcpm
  image: nullableeth/voxcpm-wyoming:latest
  restart: unless-stopped
  ports:
    - "10211:10300"
  environment:
    VOXCPM_VOICES: "Dave:A calm male voice, professional tone|Ella:A warm southern female voice, friendly and clear"
    VOXCPM_DEVICE: "cuda"
    VOXCPM_OPTIMIZE: "true"
    VOXCPM_CFG: "2.0"
    VOXCPM_TIMESTEPS: "6"
    VOXCPM_RETRY_BADCASE: "true"
    VOXCPM_REFERENCE_WAV: ""
    VOXCPM_PROMPT_WAV: ""
    VOXCPM_PROMPT_TEXT: ""
  volumes:
    - /path/to/voice-refs:/voice-refs:ro
  deploy:
    resources:
      reservations:
        devices:
          - driver: nvidia
            count: 1
            capabilities: [gpu]

⁠Implementation notes

  • Inference runs on the main asyncio thread (not run_in_executor) — required for torch.compile CUDA graph consistency
  • torch.set_float32_matmul_precision('high') enabled for free TF32 tensor core utilization
  • TORCHINDUCTOR_DISABLE_CUDAGRAPH=1 is not set — CUDA graphs are active and working correctly
  • Model weights baked into image at build time via huggingface_hub.snapshot_download
  • local_files_only=True at runtime — no network access after image build

ImagePurposeVRAMLatency
nullableeth/whisperx-wyomingSpeech-to-text (WhisperX)~1-2GB—
nullableeth/kokoro-wyomingTTS (Kokoro ONNX, fast)~720MB<0.5s
nullableeth/rvc-wyomingVoice conversion (RVC)~1GB~0.5s
nullableeth/voxcpm-wyomingTTS (VoxCPM2, high quality)~5.9GB~1.3s+

Tag summary

Content type

Image

Digest

sha256:c12be8f77…

Size

7.2 GB

Last updated

6 months ago

docker pull nullableeth/voxcpm-wyoming