StyleTTS2 voice cloning TTS with Wyoming protocol. Fast, high-quality, natural prosody for HA.
1.2K
High-quality, expressive Text-to-Speech using StyleTTS2 with style and prosody preservation. Wyoming protocol integration for seamless Home Assistant compatibility. Fast synthesis with excellent voice quality.
services:
styletts2-wyoming:
container_name: styletts2-wyoming
image: nullableeth/styletts2-wyoming:latest
restart: unless-stopped
ports:
- 10700:10700
environment:
SPEAKERS_DIR: "/app/speakers"
MODEL_NAME: "LibriTTS" # or "LJSpeech"
volumes:
- /path/to/your/speakers:/app/speakers
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
StyleTTS2 is an end-to-end text-to-speech model that uses style transfer and diffusion:
Process flow:
Text → Phonemization → Style Encoder → Diffusion Model → High-Quality Audio → Wyoming → Home Assistant
Key advantages:
StyleTTS2 only requires a reference audio sample - no training or embeddings needed. Voice characteristics are computed on-the-fly.
Requirements:
# Create speaker directory
mkdir -p /path/to/speakers/my_voice
# Copy reference audio
cp /path/to/recording.wav /path/to/speakers/my_voice/reference.wav
# Restart container
docker restart styletts2-wyoming
That's it! StyleTTS2 automatically computes style embeddings from the reference audio on startup.
speakers/
john_doe/
reference.wav # 10-30 seconds of natural speech
jane_smith/
reference.wav # Only reference.wav needed per voice
my_voice/
reference.wav
For best results:
Good reference audio examples:
Poor reference audio:
Style Preservation: StyleTTS2 captures speaking style, rhythm, and prosody from your reference. More expressive reference audio = more expressive synthesis.
| Variable | Default | Description |
|---|---|---|
SPEAKERS_DIR | /app/speakers | Directory containing voice references |
MODEL_NAME | LibriTTS | Base model to use: LibriTTS or LJSpeech |
StyleTTS2 includes two pre-trained base models. Choose the one that best matches your voice type:
When to try LJSpeech:
Set via environment variable:
environment:
MODEL_NAME: "LJSpeech"
Both models are pre-loaded in the image - no downloads needed, just change the environment variable and restart.
The container runs a Wyoming protocol server that Home Assistant can discover and use.
your-server-ip10700action:
- service: tts.speak
target:
entity_id: media_player.living_room
data:
message: "Hello! This is StyleTTS2 with natural prosody."
media_player_entity_id: media_player.living_room
options:
voice: my_voice # Your custom voice name
Works with Home Assistant's Assist voice assistant:
Voice styles are precomputed and cached on startup for instant synthesis:
2026-03-18 05:41:36 - INFO - Loading StyleTTS2 model: LibriTTS
2026-03-18 05:41:43 - INFO - Model preloaded and ready
2026-03-18 05:41:43 - INFO - Precomputing style embeddings...
2026-03-18 05:42:02 - INFO - Cached style for: john_doe
2026-03-18 05:42:02 - INFO - Cached style for: jane_smith
2026-03-18 05:42:02 - INFO - Cached 2 voice styles
Long text is automatically split into natural chunks (max 200 characters):
2026-03-17 23:01:00 - INFO - Split into 16 chunks
Unnecessary warnings suppressed for clean logs - only important messages shown.
No voices available
docker exec styletts2-wyoming ls /app/speakersreference.wav exists in each voice folderdocker logs styletts2-wyomingVoice doesn't preserve style
MODEL_NAME to LJSpeech if using LibriTTS (or vice versa)Slow synthesis
docker exec styletts2-wyoming nvidia-smiAudio cuts off or errors on long text
Poor audio quality
Container won't start
docker logs styletts2-wyomingMODEL_NAME is set to LibriTTS or LJSpeech# Add several voices at once
for voice in alice bob charlie; do
mkdir -p /path/to/speakers/$voice
cp /recordings/${voice}.wav /path/to/speakers/$voice/reference.wav
done
docker restart styletts2-wyoming
# Replace reference audio (style will be recomputed on restart)
cp new_recording.wav /path/to/speakers/john_doe/reference.wav
docker restart styletts2-wyoming
# Try LJSpeech model instead of LibriTTS
# Edit your docker-compose.yml or set environment variable:
docker stop styletts2-wyoming
# Update MODEL_NAME: "LJSpeech" in docker-compose.yml
docker start styletts2-wyoming
# Check which model is loaded in logs:
docker logs styletts2-wyoming | grep "Loading StyleTTS2 model"
# Watch logs in real-time
docker logs -f styletts2-wyoming
# Check available voices
docker exec styletts2-wyoming ls /app/speakers
# Verify GPU usage during synthesis
docker exec styletts2-wyoming nvidia-smi
| Feature | StyleTTS2 | XTTS | Piper+So-VITS |
|---|---|---|---|
| Speed | ⚡⚡⚡ Fast | ⚡⚡ Medium | ⚡ Slow |
| Quality | ⭐⭐⭐ Excellent | ⭐⭐⭐ Excellent | ⭐⭐ Good |
| Prosody | ⭐⭐⭐ Natural | ⭐⭐ Good | ⭐ Robotic |
| Setup | ✅ Simple | ⚠️ Embedding gen | ⚠️ Model training |
| VRAM | 3GB | 4GB | 4GB |
StyleTTS2 model is available under MIT License. See StyleTTS2 License for details.
Content type
Image
Digest
sha256:da66ee896…
Size
5.5 GB
Last updated
7 months ago
docker pull nullableeth/styletts2-wyoming