Sign inSign up

nullableeth/xtts-wyoming

By nullableeth

•Updated 7 months ago

XTTS voice cloning TTS for Home Assistant. Wyoming protocol. Natural speech, multi-voice support.

Image
Machine learning & AI
0

1.9K

nullableeth/xtts-wyoming repository overview

⁠TLDR: Second attempt at voice synthesis/changing/cloning/custom voice, much better a few seconds on 4070, audio considerably better than previous.

⁠XTTS Wyoming - Multi-Voice Text-to-Speech

High-quality, natural-sounding Text-to-Speech using Coqui XTTS v2 with custom voice cloning. Wyoming protocol integration for seamless Home Assistant compatibility.

⁠Features

  • 🎯 High-quality synthesis - Natural, expressive voice cloning
  • ⚡ Configurable speed - Adjust speech rate and add silence padding
  • 🎤 Multiple voices - Support for multiple custom voice profiles
  • 🏠 Home Assistant ready - Wyoming protocol integration
  • 🔧 Easy voice addition - Simple script to add new voices
  • 🚀 GPU accelerated - Fast inference with NVIDIA GPU

⁠Quick Start

⁠Docker Compose (Home Assistant)
services:
  xtts-wyoming:
    container_name: xtts-wyoming
    image: nullableeth/xtts-wyoming:latest
    restart: unless-stopped
    ports:
      - 10600:10600
    environment:
      COQUI_TOS_AGREED: "1"
      XTTS_HOST: "localhost"
      XTTS_PORT: "80"
      SPEAKERS_DIR: "/app/speakers"
      XTTS_SPEED: "1.0"           # Speech speed (0.5-2.0)
      SILENCE_PAD_MS: "200"       # Silence padding in milliseconds
      STREAM_CHUNK_SIZE: "20"     # Lower = faster streaming (5-20)
    volumes:
      - /path/to/your/speakers:/app/speakers
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

⁠Voice Setup

⁠Adding Custom Voices

XTTS requires a reference audio sample to clone a voice. The add_voice.sh script handles embedding generation automatically.

Requirements:

  • Reference audio: 6-30 seconds of clean speech
  • Format: WAV file (any sample rate, mono or stereo)
  • Quality: Clear recording, minimal background noise
  • Content: Natural speaking (not singing), preferably varied intonation

⁠Adding New Voices with Pre-computed Embeddings

For faster first-time inference, you can pre-compute voice embeddings. This helper script generates embeddings before the container needs them.

⁠Helper Script: add_voice.sh

Save this script on your host machine:

#!/bin/bash
# Usage: ./add_voice.sh speaker_name /path/to/reference.wav
set -e

SPEAKER_NAME="$1"
REFERENCE_WAV="$2"
SPEAKERS_DIR="/path/to/your/speakers"  # Change this to your speakers directory
CONTAINER_NAME="xtts"  # Change if your container has a different name

if [ -z "$SPEAKER_NAME" ]; then
    echo "Usage: $0 <speaker_name> [/path/to/reference.wav]"
    echo "Example: $0 john_doe ~/audio/john_voice.wav"
    exit 1
fi

SPEAKER_DIR="$SPEAKERS_DIR/$SPEAKER_NAME"

# Create speaker directory
mkdir -p "$SPEAKER_DIR"

# Copy reference WAV if provided
if [ -n "$REFERENCE_WAV" ]; then
    if [ ! -f "$REFERENCE_WAV" ]; then
        echo "Error: Reference file not found: $REFERENCE_WAV"
        exit 1
    fi
    
    echo "Copying reference audio..."
    cp "$REFERENCE_WAV" "$SPEAKER_DIR/reference.wav"
fi

# Check if reference.wav exists
if [ ! -f "$SPEAKER_DIR/reference.wav" ]; then
    echo "Error: No reference.wav found in $SPEAKER_DIR/"
    echo "Please provide a reference WAV file"
    exit 1
fi

echo "Generating embeddings for '$SPEAKER_NAME'..."

# Copy reference to container
docker cp "$SPEAKER_DIR/reference.wav" $CONTAINER_NAME:/tmp/reference.wav

# Generate embeddings inside container
docker exec $CONTAINER_NAME curl -s -X POST http://localhost:80/clone_speaker \
  -F "wav_file=@/tmp/reference.wav" \
  -F "speaker_name=$SPEAKER_NAME" \
  -o /tmp/embeddings.json

# Copy embeddings back
docker cp $CONTAINER_NAME:/tmp/embeddings.json "$SPEAKER_DIR/embeddings.json"

# Clean up temp files in container
docker exec $CONTAINER_NAME rm /tmp/reference.wav /tmp/embeddings.json

# Check embeddings were created
if [ ! -f "$SPEAKER_DIR/embeddings.json" ]; then
    echo "Error: embeddings.json not created"
    exit 1
fi

# Verify it's valid JSON
if ! python3 -c "import json; json.load(open('$SPEAKER_DIR/embeddings.json'))" 2>/dev/null; then
    echo "Error: Invalid embeddings file"
    cat "$SPEAKER_DIR/embeddings.json"
    exit 1
fi

EMBEDDINGS_SIZE=$(stat -f%z "$SPEAKER_DIR/embeddings.json" 2>/dev/null || stat -c%s "$SPEAKER_DIR/embeddings.json")
echo "Embeddings generated: $EMBEDDINGS_SIZE bytes"

# Restart Wyoming wrapper to load new voice
echo "Restarting $CONTAINER_NAME container..."
docker restart $CONTAINER_NAME

echo ""
echo "✅ Voice '$SPEAKER_NAME' added successfully!"
echo "   Location: $SPEAKER_DIR/"
echo "   Reference: $SPEAKER_DIR/reference.wav"
echo "   Embeddings: $SPEAKER_DIR/embeddings.json"
echo ""
echo "The voice will be available after the container restarts (~90 seconds)"

Usage:

# 1. Save script and make executable
chmod +x add_voice.sh

# 2. Edit SPEAKERS_DIR and CONTAINER_NAME at top of script

# 3. Add a voice with pre-computed embeddings
./add_voice.sh alice ~/audio/alice_sample.wav

# Wait ~90 seconds for container restart, then voice is ready!

Benefits of pre-computing embeddings:

  • ✅ First inference is instant (no 3-5 second delay)
  • ✅ Embeddings are cached and reused
  • ✅ Faster voice switching in Home Assistant
  • ✅ No compute overhead on first TTS request

Without embeddings (manual method):

  • Container computes embeddings on first use
  • 3-5 second delay on first synthesis
  • Subsequent uses are fast (cached)

What the script does:

  1. Copies reference audio to speaker directory
  2. Connects to XTTS server inside container
  3. Generates voice embeddings automatically
  4. Saves embeddings.json alongside reference audio
  5. Restarts Wyoming wrapper to load new voice
⁠Manual Voice Setup

Alternatively, prepare voices manually:

# Create speaker directory
mkdir -p /path/to/speakers/my_voice

# Copy reference audio (6-30 seconds of clean speech)
cp /path/to/recording.wav /path/to/speakers/my_voice/reference.wav

# Generate embeddings using the script
docker exec xtts-wyoming /app/add_voice.sh "my_voice" /app/speakers/my_voice/reference.wav

# Restart container
docker restart xtts-wyoming
⁠Directory Structure
speakers/
  john_doe/
    reference.wav       # 6-30 seconds of clean speech
    embeddings.json     # Auto-generated voice embeddings
  jane_smith/
    reference.wav
    embeddings.json
⁠Voice Recording Best Practices

For best results:

  • Duration: 10-15 seconds ideal (minimum 6 seconds, maximum 30 seconds)
  • Content: Natural conversational speech with varied pitch and tone
  • Avoid: Singing, monotone reading, background music
  • Quality: Use decent microphone in quiet room
  • Format: WAV preferred, but MP3/M4A work (converted automatically)

Good reference audio examples:

  • "Hey there! I'm recording this sample for voice cloning. It's important to speak naturally and clearly, with some variation in my tone and emotion."
  • Excerpt from natural conversation or podcast

Poor reference audio:

  • Singing or humming
  • Robotic/monotone reading
  • Music in background
  • Very quiet or distant recording

⁠Configuration

⁠Environment Variables
VariableDefaultDescription
COQUI_TOS_AGREED1Accept Coqui TOS (required)
XTTS_HOSTlocalhostXTTS server host
XTTS_PORT80XTTS server port
SPEAKERS_DIR/app/speakersDirectory containing voice references
XTTS_SPEED1.0Speech speed (0.5 = slow, 2.0 = fast)
SILENCE_PAD_MS0Silence padding before/after speech (milliseconds)
STREAM_CHUNK_SIZE20Streaming chunk size (lower = faster start, 5-20)
⁠Speed Adjustment
  • XTTS_SPEED: "0.8" - 20% slower (more deliberate)
  • XTTS_SPEED: "1.0" - Normal speed (default)
  • XTTS_SPEED: "1.2" - 20% faster (more energetic)
⁠Silence Padding

Add pauses before and after speech for more natural delivery:

  • SILENCE_PAD_MS: "0" - No padding (default)
  • SILENCE_PAD_MS: "200" - 200ms pause before/after
  • SILENCE_PAD_MS: "500" - 500ms pause (for dramatic effect)

⁠Home Assistant Integration

  1. Settings → Devices & Services → Add Integration
  2. Search for "Wyoming Protocol"
  3. Host: your-server-ip, Port: 10600
  4. Select voice from your speakers directory
⁠Using in Automations
action:
  - service: tts.speak
    target:
      entity_id: media_player.living_room
    data:
      message: "Hello! This is my cloned voice speaking."
      media_player_entity_id: media_player.living_room
      options:
        voice: john_doe  # Your custom voice name

⁠Performance

  • First request: ~5-10 seconds (model loading)
  • Subsequent requests: ~4-8 seconds per paragraph
  • GPU: Required for acceptable performance
  • VRAM usage: ~3-4GB
  • Quality: Excellent voice cloning, natural prosody

⁠Troubleshooting

No voices available

  • Check speaker directory mount: docker exec xtts-wyoming ls /app/speakers
  • Verify reference.wav and embeddings.json exist in each voice folder
  • Check logs: docker logs xtts-wyoming

Voice doesn't sound like reference

  • Use longer reference audio (15-20 seconds)
  • Ensure reference has varied intonation (not monotone)
  • Try different reference audio with clearer speech
  • Avoid music or background noise in reference

Slow synthesis

  • Reduce STREAM_CHUNK_SIZE to 10 for faster streaming
  • Verify GPU access: docker exec xtts-wyoming nvidia-smi
  • Check VRAM isn't exhausted (need ~4GB free)

Embeddings generation fails

  • Ensure reference.wav is valid audio file
  • Try converting to WAV: ffmpeg -i input.mp3 -ar 22050 output.wav
  • Check container has enough disk space

Voice quality issues

  • Increase XTTS_SPEED slightly (try 1.1)
  • Add silence padding: SILENCE_PAD_MS: "200"
  • Use higher quality reference audio

⁠Advanced Usage

⁠Adding Multiple Voices at Once
# Add several voices
for voice in alice bob charlie; do
  docker exec xtts-wyoming /app/add_voice.sh "$voice" "/path/to/${voice}_voice.wav"
done
⁠Updating Existing Voice
# Replace reference audio and regenerate embeddings
cp new_recording.wav /path/to/speakers/john_doe/reference.wav
docker exec xtts-wyoming /app/add_voice.sh "john_doe" /app/speakers/john_doe/reference.wav
⁠Monitoring Synthesis
# Watch logs in real-time
docker logs -f xtts-wyoming

# Check available voices
docker exec xtts-wyoming ls /app/speakers

⁠Technical Details

  • Model: Coqui XTTS v2 (multilingual)
  • Embedding: Automatic generation via XTTS API
  • Protocol: Wyoming for Home Assistant integration
  • Streaming: Chunked audio delivery for faster perceived response
  • Languages: Supports multiple languages (English, Spanish, French, etc.)

⁠Limitations

  • Accent preservation: May not fully preserve strong regional accents
  • Singing: Not designed for singing voice cloning
  • Speed: Slower than simpler TTS engines like Piper
  • VRAM: Requires GPU with 4GB+ VRAM

⁠License

XTTS model is licensed under Coqui Public Model License. See Coqui License⁠ for details.

Tag summary

Content type

Image

Digest

sha256:3de1220f1…

Size

9.6 GB

Last updated

7 months ago

docker pull nullableeth/xtts-wyoming