TTS with kokoro onnx. Fully GPU accelerated. Wyoming protocol for Home Assistant.
4.8K
High-performance GPU-accelerated Kokoro TTS server with Wyoming protocol support for Home Assistant and other applications.
Existing Kokoro Wyoming containers run on CPU, resulting in generation times of 2-4 seconds per sentence. This container leverages CUDA GPU acceleration to achieve sub-second generation times (~0.2-0.6s), making it suitable for real-time voice applications.
| Implementation | Hardware | Generation Time* | Real-time Factor |
|---|---|---|---|
| CPU-only containers | Any CPU | 2-4 seconds | 0.5-1.0x |
| This container (GPU) | NVIDIA GPU | 0.2-0.6 seconds | ~10-30x |
*For a typical 5-second audio output
docker run -d \
--name kokoro-tts \
--gpus all \
-p 10210:10210 \
nullableeth/kokoro-wyoming:latest
docker run -d \
--name kokoro-tts \
--gpus all \
-e KOKORO_SPEED=1.3 \
-p 10210:10210 \
nullableeth/kokoro-wyoming:latest
services:
kokoro-tts:
container_name: kokoro-tts
image: nullableeth/kokoro-wyoming:latest
restart: unless-stopped
ports:
- "10210:10210"
environment:
KOKORO_SPEED: "1.0" # 0.5 - 2.0
KOKORO_QUANTIZATION: "fp32" # fp32 recommended for GPU
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
| Variable | Default | Options | Description |
|---|---|---|---|
WYOMING_PORT | 10210 | Any port | Wyoming server port |
KOKORO_SPEED | 1.0 | 0.5 - 2.0 | Speech speed multiplier |
KOKORO_QUANTIZATION | fp32 | fp32, fp16, int8 | Model quantization level |
IMPORTANT: Benchmarking shows that on NVIDIA GPUs, fp32 is fastest. The quantized models (fp16, int8) are optimized for CPU inference and run significantly slower on GPU due to excessive CPU↔GPU memory transfers.
| Quantization | Model Size | VRAM | GPU Performance | CPU Performance | Recommendation |
|---|---|---|---|---|---|
| fp32 | 310 MB | ~2.5 GB | Fastest (0.4s) ✓ | Slow | GPU: Use this |
| fp16 | 169 MB | ~1.5 GB | 4x slower (1.7s) | Medium | Not recommended |
| int8 | 88 MB | ~1 GB | 26x slower (11s) | Fast | CPU only |
Benchmark Results (RTX 4070, 5-second audio):
Why fp32 is fastest on GPU:
When to use int8:
For GPU inference, always use fp32.
0.8 - Slower, more deliberate speech1.0 - Normal speech rate1.3 - Faster, more energetic speech1.5 - Very fast speechNote: Speed affects generation time proportionally. 2.0x speed = ~50% faster generation.
The container includes 54 voices from Kokoro v1.0:
af_alloy, af_aoede, af_bella, af_heart, af_jessica, af_kore, af_nicole, af_nova, af_river, af_sarah, af_skyam_adam, am_echo, am_eric, am_fenrir, am_liam, am_michael, am_onyx, am_puck, am_santabf_alice, bf_emma, bf_isabella, bf_lilybm_daniel, bm_fable, bm_george, bm_lewisef_doraem_alex, em_santaff_siwishf_alpha, hf_betahm_omega, hm_psiif_saraim_nicolajf_alpha, jf_gongitsune, jf_nezumi, jf_tebukurojm_kumopf_dorapm_alex, pm_santazf_xiaobei, zf_xiaoni, zf_xiaoxiao, zf_xiaoyizm_yunjian, zm_yunxi, zm_yunxia, zm_yunyangVoice Naming Convention:
- First letter: Language (
a=American,b=British,e=Spanish,f=French,h=Hindi,i=Italian,j=Japanese,p=Portuguese,z=Chinese)- Second letter: Gender (
f=Female,m=Male)- Rest: Voice name
tts:
- platform: wyoming
host: 192.168.1.100 # Your Docker host IP
port: 10210
voice: af_sky # Default voice
service: tts.speak
data:
entity_id: media_player.living_room
message: "Hello from Kokoro TTS"
options:
voice: af_bella # Choose any of the 54 voices
GPU (RTX 4070) - Always use fp32:
| Speed | Generation Time | VRAM Usage | Use Case |
|---|---|---|---|
| 1.0x | 0.44s | 2.5 GB | Default, best quality |
| 1.3x | 0.34s | 2.5 GB | Faster speech |
| 1.5x | 0.29s | 2.5 GB | Very fast speech |
| 2.0x | 0.22s | 2.5 GB | Maximum speed |
For ~5 second audio output with fp32 quantization
CPU Mode (use int8):
environment:
KOKORO_QUANTIZATION: "int8" # Optimized for CPU
# Remove the deploy.resources.reservations section
Using Wyoming client tools:
# Install wyoming client
pip install wyoming
# Generate speech
echo "Hello world" | wyoming-client \
--host localhost \
--port 10210 \
--voice af_sky \
--output output.wav
Kokoro can be chained with RVC (Retrieval-based Voice Conversion) for custom voices:
services:
kokoro-tts:
image: nullableeth/kokoro-wyoming:latest
ports:
- "10210:10210"
environment:
KOKORO_QUANTIZATION: "fp32" # Fast GPU inference
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
rvc-converter:
image: nullableeth/rvc-wyoming:latest
ports:
- "10900:10900"
environment:
TTS_HOST: "kokoro-tts"
TTS_PORT: "10210"
TTS_VOICE: "af_sky"
MODEL_NAME: "your_model.pth"
volumes:
- ./rvc_models:/models:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
Point Home Assistant or clients to port 10900 to use RVC-converted voices.
Check if GPU is accessible:
docker exec kokoro-tts python3 -c "import torch; print('CUDA available:', torch.cuda.is_available())"
Should output: CUDA available: True
docker logs kokoro-tts | grep "Loading Kokoro"
Loading Kokoro model (fp32, speed: 1.0x)...docker logs kokoro-tts | grep Memcpy
docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi
This is expected behavior. The quantized models are optimized for CPU inference and run slower on GPU due to excessive memory transfers. Always use fp32 for GPU.
Container logs show: Invalid KOKORO_QUANTIZATION 'xyz'
Solution: Use only fp32, fp16, or int8
Container fails with: KOKORO_SPEED must be between 0.5 and 2.0
Solution: Set KOKORO_SPEED between 0.5 and 2.0
Verify port mapping and firewall settings:
docker logs kokoro-tts | grep "Starting server"
pytorch/pytorch:2.10.0-cuda12.8-cudnn9-runtimeThe container uses ONNX Runtime with CUDAExecutionProvider for GPU acceleration. Required CUDA libraries are bundled in the image via pip packages (nvidia-cublas, nvidia-cudnn, etc.).
Quantization Performance Note: The int8 and fp16 models generate excessive CPU↔GPU memory transfers (550+ Memcpy operations vs 39 for fp32), making them significantly slower on GPU. These quantized models are optimized for CPU inference with AVX-512 instructions, not CUDA.
git clone <your-repo>
cd kokoro-wyoming
docker build -t kokoro-wyoming:latest .
This container packages:
For issues specific to this container, please open an issue on GitHub. For Kokoro model issues, see the official Kokoro repository.
Content type
Image
Digest
sha256:fcc58da35…
Size
5 GB
Last updated
7 months ago
docker pull nullableeth/kokoro-wyoming