XTTS voice cloning TTS for Home Assistant. Wyoming protocol. Natural speech, multi-voice support.
1.9K
High-quality, natural-sounding Text-to-Speech using Coqui XTTS v2 with custom voice cloning. Wyoming protocol integration for seamless Home Assistant compatibility.
services:
xtts-wyoming:
container_name: xtts-wyoming
image: nullableeth/xtts-wyoming:latest
restart: unless-stopped
ports:
- 10600:10600
environment:
COQUI_TOS_AGREED: "1"
XTTS_HOST: "localhost"
XTTS_PORT: "80"
SPEAKERS_DIR: "/app/speakers"
XTTS_SPEED: "1.0" # Speech speed (0.5-2.0)
SILENCE_PAD_MS: "200" # Silence padding in milliseconds
STREAM_CHUNK_SIZE: "20" # Lower = faster streaming (5-20)
volumes:
- /path/to/your/speakers:/app/speakers
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
XTTS requires a reference audio sample to clone a voice. The add_voice.sh script handles embedding generation automatically.
Requirements:
For faster first-time inference, you can pre-compute voice embeddings. This helper script generates embeddings before the container needs them.
add_voice.shSave this script on your host machine:
#!/bin/bash
# Usage: ./add_voice.sh speaker_name /path/to/reference.wav
set -e
SPEAKER_NAME="$1"
REFERENCE_WAV="$2"
SPEAKERS_DIR="/path/to/your/speakers" # Change this to your speakers directory
CONTAINER_NAME="xtts" # Change if your container has a different name
if [ -z "$SPEAKER_NAME" ]; then
echo "Usage: $0 <speaker_name> [/path/to/reference.wav]"
echo "Example: $0 john_doe ~/audio/john_voice.wav"
exit 1
fi
SPEAKER_DIR="$SPEAKERS_DIR/$SPEAKER_NAME"
# Create speaker directory
mkdir -p "$SPEAKER_DIR"
# Copy reference WAV if provided
if [ -n "$REFERENCE_WAV" ]; then
if [ ! -f "$REFERENCE_WAV" ]; then
echo "Error: Reference file not found: $REFERENCE_WAV"
exit 1
fi
echo "Copying reference audio..."
cp "$REFERENCE_WAV" "$SPEAKER_DIR/reference.wav"
fi
# Check if reference.wav exists
if [ ! -f "$SPEAKER_DIR/reference.wav" ]; then
echo "Error: No reference.wav found in $SPEAKER_DIR/"
echo "Please provide a reference WAV file"
exit 1
fi
echo "Generating embeddings for '$SPEAKER_NAME'..."
# Copy reference to container
docker cp "$SPEAKER_DIR/reference.wav" $CONTAINER_NAME:/tmp/reference.wav
# Generate embeddings inside container
docker exec $CONTAINER_NAME curl -s -X POST http://localhost:80/clone_speaker \
-F "wav_file=@/tmp/reference.wav" \
-F "speaker_name=$SPEAKER_NAME" \
-o /tmp/embeddings.json
# Copy embeddings back
docker cp $CONTAINER_NAME:/tmp/embeddings.json "$SPEAKER_DIR/embeddings.json"
# Clean up temp files in container
docker exec $CONTAINER_NAME rm /tmp/reference.wav /tmp/embeddings.json
# Check embeddings were created
if [ ! -f "$SPEAKER_DIR/embeddings.json" ]; then
echo "Error: embeddings.json not created"
exit 1
fi
# Verify it's valid JSON
if ! python3 -c "import json; json.load(open('$SPEAKER_DIR/embeddings.json'))" 2>/dev/null; then
echo "Error: Invalid embeddings file"
cat "$SPEAKER_DIR/embeddings.json"
exit 1
fi
EMBEDDINGS_SIZE=$(stat -f%z "$SPEAKER_DIR/embeddings.json" 2>/dev/null || stat -c%s "$SPEAKER_DIR/embeddings.json")
echo "Embeddings generated: $EMBEDDINGS_SIZE bytes"
# Restart Wyoming wrapper to load new voice
echo "Restarting $CONTAINER_NAME container..."
docker restart $CONTAINER_NAME
echo ""
echo "✅ Voice '$SPEAKER_NAME' added successfully!"
echo " Location: $SPEAKER_DIR/"
echo " Reference: $SPEAKER_DIR/reference.wav"
echo " Embeddings: $SPEAKER_DIR/embeddings.json"
echo ""
echo "The voice will be available after the container restarts (~90 seconds)"
Usage:
# 1. Save script and make executable
chmod +x add_voice.sh
# 2. Edit SPEAKERS_DIR and CONTAINER_NAME at top of script
# 3. Add a voice with pre-computed embeddings
./add_voice.sh alice ~/audio/alice_sample.wav
# Wait ~90 seconds for container restart, then voice is ready!
Benefits of pre-computing embeddings:
Without embeddings (manual method):
What the script does:
embeddings.json alongside reference audioAlternatively, prepare voices manually:
# Create speaker directory
mkdir -p /path/to/speakers/my_voice
# Copy reference audio (6-30 seconds of clean speech)
cp /path/to/recording.wav /path/to/speakers/my_voice/reference.wav
# Generate embeddings using the script
docker exec xtts-wyoming /app/add_voice.sh "my_voice" /app/speakers/my_voice/reference.wav
# Restart container
docker restart xtts-wyoming
speakers/
john_doe/
reference.wav # 6-30 seconds of clean speech
embeddings.json # Auto-generated voice embeddings
jane_smith/
reference.wav
embeddings.json
For best results:
Good reference audio examples:
Poor reference audio:
| Variable | Default | Description |
|---|---|---|
COQUI_TOS_AGREED | 1 | Accept Coqui TOS (required) |
XTTS_HOST | localhost | XTTS server host |
XTTS_PORT | 80 | XTTS server port |
SPEAKERS_DIR | /app/speakers | Directory containing voice references |
XTTS_SPEED | 1.0 | Speech speed (0.5 = slow, 2.0 = fast) |
SILENCE_PAD_MS | 0 | Silence padding before/after speech (milliseconds) |
STREAM_CHUNK_SIZE | 20 | Streaming chunk size (lower = faster start, 5-20) |
XTTS_SPEED: "0.8" - 20% slower (more deliberate)XTTS_SPEED: "1.0" - Normal speed (default)XTTS_SPEED: "1.2" - 20% faster (more energetic)Add pauses before and after speech for more natural delivery:
SILENCE_PAD_MS: "0" - No padding (default)SILENCE_PAD_MS: "200" - 200ms pause before/afterSILENCE_PAD_MS: "500" - 500ms pause (for dramatic effect)your-server-ip, Port: 10600action:
- service: tts.speak
target:
entity_id: media_player.living_room
data:
message: "Hello! This is my cloned voice speaking."
media_player_entity_id: media_player.living_room
options:
voice: john_doe # Your custom voice name
No voices available
docker exec xtts-wyoming ls /app/speakersreference.wav and embeddings.json exist in each voice folderdocker logs xtts-wyomingVoice doesn't sound like reference
Slow synthesis
STREAM_CHUNK_SIZE to 10 for faster streamingdocker exec xtts-wyoming nvidia-smiEmbeddings generation fails
ffmpeg -i input.mp3 -ar 22050 output.wavVoice quality issues
XTTS_SPEED slightly (try 1.1)SILENCE_PAD_MS: "200"# Add several voices
for voice in alice bob charlie; do
docker exec xtts-wyoming /app/add_voice.sh "$voice" "/path/to/${voice}_voice.wav"
done
# Replace reference audio and regenerate embeddings
cp new_recording.wav /path/to/speakers/john_doe/reference.wav
docker exec xtts-wyoming /app/add_voice.sh "john_doe" /app/speakers/john_doe/reference.wav
# Watch logs in real-time
docker logs -f xtts-wyoming
# Check available voices
docker exec xtts-wyoming ls /app/speakers
XTTS model is licensed under Coqui Public Model License. See Coqui License for details.
Content type
Image
Digest
sha256:3de1220f1…
Size
9.6 GB
Last updated
7 months ago
docker pull nullableeth/xtts-wyoming