Zero-shot voice cloning TTS with Qwen3-TTS. Wyoming protocol for Home Assistant. GPU-accelerated.
1.9K
GPU-accelerated Qwen3-TTS server with Wyoming protocol support for Home Assistant and other applications.
This container provides access to all three Qwen3-TTS model types:
Use pre-trained voices with high quality.
services:
wyoming-qwentts:
container_name: wyoming-qwentts
image: nullableeth/qwentts-wyoming:latest
restart: unless-stopped
ports:
- "10204:10800"
environment:
QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice"
QWEN_VOICE: "ryan" # serena, vivian, uncle_fu, ryan, aiden, ono_anna, sohee, eric, dylan
QWEN_LANGUAGE: "english"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
Available voices: serena, vivian, uncle_fu, ryan, aiden, ono_anna, sohee, eric, dylan
Generate voices from natural language descriptions.
services:
wyoming-qwentts:
container_name: wyoming-qwentts
image: nullableeth/qwentts-wyoming:latest
restart: unless-stopped
ports:
- "10204:10800"
environment:
QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign"
VOICE_DESIGN_DESC: "A female voice with a soft and gentle tone"
QWEN_LANGUAGE: "english"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
Note: VoiceDesign is only available in the 1.7B model.
Clone voices from 3-second audio reference samples.
services:
wyoming-qwentts:
container_name: wyoming-qwentts
image: nullableeth/qwentts-wyoming:latest
restart: unless-stopped
ports:
- "10204:10800"
environment:
QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-Base"
QWEN_LANGUAGE: "english"
volumes:
- ./voices:/app/voices:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
Voice directory structure:
voices/
├── ella/
│ ├── reference.wav (3+ seconds of clean audio)
│ └── reference.txt (transcript of the audio)
└── morgan/
├── reference.wav
└── reference.txt
Each directory name becomes a selectable voice in Wyoming (e.g., "ella", "morgan").
| Variable | Default | Options | Description |
|---|---|---|---|
QWEN_MODEL | Qwen/Qwen3-TTS-12Hz-0.6B-Base | See models below | Model to load |
QWEN_LANGUAGE | english | See languages below | Synthesis language |
QWEN_VOICE | Auto-select | Speaker name | Voice for CustomVoice models |
VOICE_DESIGN_DESC | Default description | Any text | Voice description for VoiceDesign |
| Model | Size | Type | Features |
|---|---|---|---|
Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice | 0.6B | CustomVoice | 9 premium voices, faster |
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice | 1.7B | CustomVoice | 9 premium voices, higher quality |
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign | 1.7B | VoiceDesign | Generate from description |
Qwen/Qwen3-TTS-12Hz-0.6B-Base | 0.6B | Base | Voice cloning, faster |
Qwen/Qwen3-TTS-12Hz-1.7B-Base | 1.7B | Base | Voice cloning, higher quality |
auto - Auto-detectenglishchinesejapanesekoreangermanfrenchrussianportuguesespanishitalianMale voices: ryan, aiden, eric, dylan, uncle_fu
Female voices: serena, vivian, ono_anna, sohee
Add to configuration.yaml:
tts:
- platform: wyoming
host: 192.168.1.100 # Your Docker host IP
port: 10204
voice: ryan # For CustomVoice models
Test in Home Assistant:
service: tts.speak
data:
entity_id: media_player.living_room
message: "Hello from Qwen3 TTS"
options:
voice: serena # Choose any available voice
GPU (RTX 4070):
| Model | Voice Type | Generation Time* | VRAM Usage |
|---|---|---|---|
| 0.6B CustomVoice | ryan | ~8-12s | ~3.7 GB |
| 1.7B CustomVoice | ryan | ~8-12s | ~4.0 GB |
| 0.6B Base | cloned | ~8-12s | ~3.7 GB |
| 1.7B Base | cloned | ~8-12s | ~4.0 GB |
*For a typical 20-word sentence
Note: Qwen models are optimized for quality over speed. For faster generation (<1s), consider using Kokoro TTS with RVC voice conversion.
Using Wyoming client tools:
# Install wyoming client
pip install wyoming
# Generate speech
echo "Hello world" | wyoming-client \
--host localhost \
--port 10204 \
--voice ryan \
--output output.wav
For best results, use Base model for prosody, then apply RVC for custom voice:
services:
qwen-tts:
image: nullableeth/qwentts-wyoming:latest
ports:
- "10204:10800"
environment:
QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-Base"
volumes:
- ./voices:/app/voices:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
rvc-converter:
image: nullableeth/rvc-wyoming:latest
ports:
- "10900:10900"
environment:
TTS_HOST: "qwen-tts"
TTS_PORT: "10800"
TTS_VOICE: "ella" # Your cloned voice
MODEL_NAME: "your_model.pth"
volumes:
- ./rvc_models:/models:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
Point Home Assistant to port 10900 for RVC-converted output.
Cause: Using a Base model without mounting /app/voices.
Solution: Either:
/app/voices, ORCause: CustomVoice model failed to load speaker list.
Solution: Check container logs:
docker logs wyoming-qwentts | grep -i speaker
Cause: NVIDIA Container Toolkit not installed or GPU not accessible.
Solution:
docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi
Cause: Running on CPU instead of GPU.
Solution: Ensure GPU is accessible and check logs for "Using GPU with flash-attention" message.
Cause: CustomVoice speakers may have non-native accents.
Solution: Try different speakers (ryan, aiden, serena, vivian) or use Base model with native English reference audio.
pytorch/pytorch:2.5.1-cuda12.4-cudnn9-develBase models use a 3-second reference audio sample to clone voice characteristics:
reference.wav (3+ seconds) and reference.txt (transcript)/app/voices/ becomes a selectable voiceModels are downloaded on first run from HuggingFace:
Models are cached in the container at /root/.cache/huggingface/.
git clone <your-repo>
cd qwentts-wyoming/build
docker build -t qwentts-wyoming:latest .
This container packages:
For issues specific to this container, please open an issue on GitHub. For Qwen3-TTS model issues, see the official Qwen repository.
| Model | Speed | Quality | Voice Options | Best For |
|---|---|---|---|---|
| Kokoro | ⚡ Very Fast (0.2-0.6s) | High | 54 voices | Real-time applications |
| Qwen3-TTS | Slow (8-12s) | Very High | 9 voices + cloning | Quality over speed |
| Piper | Fast (1-2s) | Medium | 100+ voices | Resource-constrained |
| StyleTTS2 | Slow (10-20s) | Very High | Voice cloning | Studio quality |
Recommendation: For sub-second generation with custom voices, use Kokoro + RVC instead of Qwen3-TTS.
Content type
Image
Digest
sha256:81477b29c…
Size
7.4 GB
Last updated
7 months ago
docker pull nullableeth/qwentts-wyoming