Fast Piper TTS + So-VITS voice cloning. Wyoming protocol for Home Assistant. GPU-accelerated.
399
Fast, high-quality Text-to-Speech combining Piper's speed with So-VITS-SVC voice cloning for natural-sounding speech with custom voices. Wyoming protocol wrapper for seamless Home Assistant integration.
services:
piper-synthesized:
container_name: piper-synthesized
image: nullableeth/piper-synthesized:latest
restart: unless-stopped
ports:
- 10200:10200
environment:
PIPER_VOICE: "en_US-lessac-medium"
PIPER_LENGTH_SCALE: "1.0"
SILENCE_PAD_MS: "200"
SPEAKERS_DIR: "/app/speakers"
volumes:
- /path/to/your/speakers:/app/speakers
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
This container combines two technologies with Wyoming protocol wrapping:
Process flow:
Text → Piper (fast synthesis) → So-VITS-SVC (voice conversion) → Wyoming → Home Assistant
To use your own voice, you'll need to train a So-VITS-SVC model:
Requirements:
Training Process:
Record audio samples
Prepare training data
Train the model
model.pth and config.json filesDeploy to container
model.pth) typically 50-200MBPre-trained Models: You can also use pre-trained So-VITS-SVC models if available, though custom training produces best results for voice matching.
speakers/
your_voice_name/
model.pth # Trained So-VITS-SVC model (~50-200MB)
config.json # Model configuration
reference.wav # Optional: sample of target voice
mkdir -p /path/to/speakers/my_voice cp /path/to/training/output/model.pth /path/to/speakers/my_voice/
cp /path/to/training/output/config.json /path/to/speakers/my_voice/
docker restart piper-synthesizedYour voice will be available as my_voice in Home Assistant via Wyoming protocol.
| Variable | Default | Description |
|---|---|---|
PIPER_VOICE | en_US-lessac-medium | Base Piper voice for synthesis |
PIPER_LENGTH_SCALE | 1.0 | Speech speed (0.5-2.0) |
SILENCE_PAD_MS | 0 | Silence padding in milliseconds |
SPEAKERS_DIR | /app/speakers | Directory containing voice models |
en_US-lessac-medium - Clear American English (default)en_US-amy-medium - Female American Englishen_US-ryan-medium - Male American Englishen_GB-alan-medium - British EnglishThe container runs a Wyoming protocol server that Home Assistant can discover and use.
your-server-ip10200action:
- service: tts.speak
target:
entity_id: media_player.living_room
data:
message: "Hello! This is my custom cloned voice."
media_player_entity_id: media_player.living_room
options:
voice: my_voice # Your custom voice name
Works with Home Assistant's Assist voice assistant:
No voices available
docker exec piper-synthesized ls /app/speakersmodel.pth and config.json exist in each voice folderdocker logs piper-synthesizedAudio sounds robotic
PIPER_LENGTH_SCALE to 1.1 for slightly more natural cadenceSILENCE_PAD_MS: "200" for better pacingSlow synthesis
docker exec piper-synthesized nvidia-smiPoor voice quality
Wyoming connection fails
docker logs piper-synthesizedContent type
Image
Digest
sha256:1ab02ceeb…
Size
4.8 GB
Last updated
7 months ago
docker pull nullableeth/piper-sovits-wyoming