Sign inSign up

nullableeth/styletts2-wyoming

By nullableeth

•Updated 7 months ago

StyleTTS2 voice cloning TTS with Wyoming protocol. Fast, high-quality, natural prosody for HA.

Image
Machine learning & AI
0

1.2K

nullableeth/styletts2-wyoming repository overview

⁠TLDR: Third attempt at voice synthesis/changing/cloning/custom voice, much much better, voice is more natural and cloning is close to original but largely based on the base model voice which work ok for women voices but fail miserably at cloning morgan freeman for example. This one however is FAST, I think it is even faster than Piper and works really well if you want to add minor changes to a females voice.

⁠StyleTTS2 Wyoming - High-Quality Voice Cloning TTS

High-quality, expressive Text-to-Speech using StyleTTS2 with style and prosody preservation. Wyoming protocol integration for seamless Home Assistant compatibility. Fast synthesis with excellent voice quality.

⁠Features

  • 🎯 Excellent quality - Natural, expressive speech synthesis
  • ⚡ Fast synthesis - Near Piper-speed performance (~1-3 seconds)
  • 🎭 Style preservation - Better prosody and intonation than basic TTS
  • 🎤 Multiple voices - Support for unlimited custom voice profiles
  • 🏠 Home Assistant ready - Wyoming protocol integration
  • 🚀 GPU accelerated - Fast inference with NVIDIA GPU
  • 📦 Pre-loaded models - No downloads on first run
  • 🔧 Model selection - Choose base model for best voice matching

⁠Quick Start

⁠Docker Compose (Home Assistant)
services:
  styletts2-wyoming:
    container_name: styletts2-wyoming
    image: nullableeth/styletts2-wyoming:latest
    restart: unless-stopped
    ports:
      - 10700:10700
    environment:
      SPEAKERS_DIR: "/app/speakers"
      MODEL_NAME: "LibriTTS"  # or "LJSpeech"
    volumes:
      - /path/to/your/speakers:/app/speakers
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

⁠How It Works

StyleTTS2 is an end-to-end text-to-speech model that uses style transfer and diffusion:

Process flow:

Text → Phonemization → Style Encoder → Diffusion Model → High-Quality Audio → Wyoming → Home Assistant

Key advantages:

  • Preserves speaking style and prosody from reference audio
  • Natural intonation and rhythm
  • Fast inference with pre-loaded models
  • Automatic text chunking for long passages

⁠Voice Setup

⁠Adding Custom Voices

StyleTTS2 only requires a reference audio sample - no training or embeddings needed. Voice characteristics are computed on-the-fly.

Requirements:

  • Reference audio: 10-30 seconds of clean, natural speech
  • Format: WAV file (any sample rate, mono or stereo)
  • Quality: Clear recording, minimal background noise
  • Content: Natural conversational speech with varied intonation
⁠Simple Voice Setup
# Create speaker directory
mkdir -p /path/to/speakers/my_voice

# Copy reference audio
cp /path/to/recording.wav /path/to/speakers/my_voice/reference.wav

# Restart container
docker restart styletts2-wyoming

That's it! StyleTTS2 automatically computes style embeddings from the reference audio on startup.

⁠Directory Structure
speakers/
  john_doe/
    reference.wav       # 10-30 seconds of natural speech
  jane_smith/
    reference.wav       # Only reference.wav needed per voice
  my_voice/
    reference.wav
⁠Voice Recording Best Practices

For best results:

  • Duration: 15-20 seconds ideal (minimum 10 seconds)
  • Content: Natural, expressive speech with varied pitch and emotion
  • Avoid: Monotone reading, whispering, shouting
  • Quality: Use decent microphone in quiet room
  • Format: WAV preferred (MP3/M4A automatically converted)

Good reference audio examples:

  • Natural conversation excerpt with emotional range
  • Storytelling with varied intonation
  • Expressive reading with natural pauses and emphasis

Poor reference audio:

  • Robotic/monotone voice
  • Singing or humming
  • Background music or noise
  • Very quiet or clipped audio
  • Heavily processed/filtered audio

Style Preservation: StyleTTS2 captures speaking style, rhythm, and prosody from your reference. More expressive reference audio = more expressive synthesis.

⁠Configuration

⁠Environment Variables
VariableDefaultDescription
SPEAKERS_DIR/app/speakersDirectory containing voice references
MODEL_NAMELibriTTSBase model to use: LibriTTS or LJSpeech
⁠Model Selection

StyleTTS2 includes two pre-trained base models. Choose the one that best matches your voice type:

  • LibriTTS (default) - General purpose, works well for most voices
  • LJSpeech - Alternative model, may work better for certain voice characteristics

When to try LJSpeech:

  • Voice doesn't match well with LibriTTS
  • Reference voice has significantly different characteristics
  • Experimentation to find best quality

Set via environment variable:

environment:
  MODEL_NAME: "LJSpeech"

Both models are pre-loaded in the image - no downloads needed, just change the environment variable and restart.

⁠Home Assistant Integration

The container runs a Wyoming protocol server that Home Assistant can discover and use.

⁠Adding to Home Assistant
  1. Settings → Devices & Services → Add Integration
  2. Search for "Wyoming Protocol"
  3. Enter:
    • Host: your-server-ip
    • Port: 10700
  4. Select your custom voice from the dropdown
⁠Using in Automations
action:
  - service: tts.speak
    target:
      entity_id: media_player.living_room
    data:
      message: "Hello! This is StyleTTS2 with natural prosody."
      media_player_entity_id: media_player.living_room
      options:
        voice: my_voice  # Your custom voice name
⁠Using with Voice Assistants

Works with Home Assistant's Assist voice assistant:

  1. Settings → Voice Assistants → Add Assistant
  2. Select Wyoming TTS as speech output
  3. Choose your custom StyleTTS2 voice

⁠Performance

  • Model loading: ~10 seconds (on startup)
  • First request: ~1-3 seconds (style cached)
  • Subsequent requests: ~1-2 seconds per sentence
  • Long text: Automatically chunked, ~2-4 seconds per chunk
  • GPU: Required for acceptable performance
  • VRAM usage: ~3GB
  • Quality: Excellent voice cloning with natural prosody

⁠Automatic Features

⁠Style Embedding Cache

Voice styles are precomputed and cached on startup for instant synthesis:

2026-03-18 05:41:36 - INFO - Loading StyleTTS2 model: LibriTTS
2026-03-18 05:41:43 - INFO - Model preloaded and ready
2026-03-18 05:41:43 - INFO - Precomputing style embeddings...
2026-03-18 05:42:02 - INFO - Cached style for: john_doe
2026-03-18 05:42:02 - INFO - Cached style for: jane_smith
2026-03-18 05:42:02 - INFO - Cached 2 voice styles
⁠Text Chunking

Long text is automatically split into natural chunks (max 200 characters):

  • Splits on sentence boundaries (. ! ?)
  • Falls back to clause boundaries (commas)
  • Falls back to word boundaries (spaces)
  • Never splits mid-word
2026-03-17 23:01:00 - INFO - Split into 16 chunks
⁠Clean Logging

Unnecessary warnings suppressed for clean logs - only important messages shown.

⁠Troubleshooting

No voices available

  • Check speaker directory mount: docker exec styletts2-wyoming ls /app/speakers
  • Verify reference.wav exists in each voice folder
  • Check logs: docker logs styletts2-wyoming

Voice doesn't preserve style

  • Use longer reference audio (15-20 seconds minimum)
  • Ensure reference has natural, varied intonation (not monotone)
  • Try reference with more emotional expression
  • Try switching base model - change MODEL_NAME to LJSpeech if using LibriTTS (or vice versa)
  • Avoid heavily compressed or filtered audio

Slow synthesis

  • Verify GPU access: docker exec styletts2-wyoming nvidia-smi
  • Check VRAM isn't exhausted (need ~4GB free)
  • Ensure CUDA drivers are properly installed

Audio cuts off or errors on long text

  • Automatic chunking should handle this
  • Check logs for "Split into X chunks" message
  • Very long text (>1000 words) may take longer

Poor audio quality

  • Use higher quality reference audio (44.1kHz or 48kHz WAV)
  • Ensure reference is clean (no background noise)
  • Try different reference audio with clearer speech
  • Experiment with base models - some voices work better with LJSpeech vs LibriTTS

Container won't start

  • Check logs: docker logs styletts2-wyoming
  • Verify GPU is accessible from Docker
  • Ensure sufficient disk space for models (~800MB)
  • If model loading fails, verify MODEL_NAME is set to LibriTTS or LJSpeech

⁠Advanced Usage

⁠Adding Multiple Voices
# Add several voices at once
for voice in alice bob charlie; do
  mkdir -p /path/to/speakers/$voice
  cp /recordings/${voice}.wav /path/to/speakers/$voice/reference.wav
done

docker restart styletts2-wyoming
⁠Updating Existing Voice
# Replace reference audio (style will be recomputed on restart)
cp new_recording.wav /path/to/speakers/john_doe/reference.wav
docker restart styletts2-wyoming
⁠Switching Base Models
# Try LJSpeech model instead of LibriTTS
# Edit your docker-compose.yml or set environment variable:
docker stop styletts2-wyoming
# Update MODEL_NAME: "LJSpeech" in docker-compose.yml
docker start styletts2-wyoming

# Check which model is loaded in logs:
docker logs styletts2-wyoming | grep "Loading StyleTTS2 model"
⁠Monitoring Synthesis
# Watch logs in real-time
docker logs -f styletts2-wyoming

# Check available voices
docker exec styletts2-wyoming ls /app/speakers

# Verify GPU usage during synthesis
docker exec styletts2-wyoming nvidia-smi

⁠Technical Details

  • Models: StyleTTS2 LibriTTS and LJSpeech checkpoints (both pre-loaded)
  • Architecture: Diffusion-based with style encoder
  • Style Transfer: Automatic from reference audio
  • Protocol: Wyoming for Home Assistant integration
  • Streaming: Chunked audio delivery
  • Port: 10700 (Wyoming protocol)
  • Image size: ~1.5GB (includes both models)

⁠Limitations

  • Accent preservation: May not fully preserve strong regional accents
  • Singing: Not designed for singing voice cloning
  • Token limit: Automatic chunking handles long text (max ~200 chars per chunk)
  • VRAM: Requires GPU with 4GB+ VRAM
  • Reference quality: Output quality depends heavily on reference audio quality
  • Model selection: Some voices work better with one base model vs the other - experimentation may be needed

⁠Comparison to Other TTS

FeatureStyleTTS2XTTSPiper+So-VITS
Speed⚡⚡⚡ Fast⚡⚡ Medium⚡ Slow
Quality⭐⭐⭐ Excellent⭐⭐⭐ Excellent⭐⭐ Good
Prosody⭐⭐⭐ Natural⭐⭐ Good⭐ Robotic
Setup✅ Simple⚠️ Embedding gen⚠️ Model training
VRAM3GB4GB4GB

⁠License

StyleTTS2 model is available under MIT License. See StyleTTS2 License⁠ for details.

Tag summary

Content type

Image

Digest

sha256:da66ee896…

Size

5.5 GB

Last updated

7 months ago

docker pull nullableeth/styletts2-wyoming