Sign inSign up

nullableeth/piper-sovits-wyoming

By nullableeth

•Updated 7 months ago

Fast Piper TTS + So-VITS voice cloning. Wyoming protocol for Home Assistant. GPU-accelerated.

Image
Machine learning & AI
0

399

nullableeth/piper-sovits-wyoming repository overview

⁠TLDR: First attempt at voice synthesis/changing/cloning/custom voice, not very good and takes are very long time for any reasonable text length.

⁠Piper Synthesized TTS with So-VITS-SVC Voice Conversion

Fast, high-quality Text-to-Speech combining Piper's speed with So-VITS-SVC voice cloning for natural-sounding speech with custom voices. Wyoming protocol wrapper for seamless Home Assistant integration.

⁠Features

  • ⚡ Fast synthesis - Piper generates speech quickly
  • 🎭 Voice cloning - So-VITS-SVC converts to your custom voice
  • 🏠 Home Assistant ready - Wyoming protocol integration
  • 🔧 Configurable - Adjust synthesis speed and silence padding
  • 🎤 Custom voices - Use your own voice samples
  • 🚀 GPU accelerated - Fast voice conversion with NVIDIA GPU

⁠Quick Start

⁠Docker Compose (Home Assistant)
services:
  piper-synthesized:
    container_name: piper-synthesized
    image: nullableeth/piper-synthesized:latest
    restart: unless-stopped
    ports:
      - 10200:10200
    environment:
      PIPER_VOICE: "en_US-lessac-medium"
      PIPER_LENGTH_SCALE: "1.0"
      SILENCE_PAD_MS: "200"
      SPEAKERS_DIR: "/app/speakers"
    volumes:
      - /path/to/your/speakers:/app/speakers
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

⁠How It Works

This container combines two technologies with Wyoming protocol wrapping:

  1. Piper TTS - Fast neural text-to-speech synthesis
  2. So-VITS-SVC - Voice conversion to match your custom voice
  3. Wyoming Protocol - Standardized TTS interface for Home Assistant

Process flow:

Text → Piper (fast synthesis) → So-VITS-SVC (voice conversion) → Wyoming → Home Assistant

⁠Voice Training & Setup

⁠Training a So-VITS-SVC Model

To use your own voice, you'll need to train a So-VITS-SVC model:

Requirements:

  • 5-10 minutes of clean voice recordings (speaking, not singing)
  • GPU with at least 8GB VRAM for training
  • Python environment with So-VITS-SVC installed

Training Process:

  1. Record audio samples

    • Record yourself speaking naturally (5-10 minutes total)
    • Use good quality microphone in quiet environment
    • Save as WAV files (22050Hz or 44100Hz sample rate)
  2. Prepare training data

    • Split recordings into short clips (5-10 seconds each)
    • Remove silence and background noise
    • Organize in a single folder
  3. Train the model

    • Use So-VITS-SVC⁠ training scripts
    • Training takes 2-6 hours on consumer GPU
    • Produces model.pth and config.json files
  4. Deploy to container

    • Copy trained files to speakers directory
    • Model file (model.pth) typically 50-200MB

Pre-trained Models: You can also use pre-trained So-VITS-SVC models if available, though custom training produces best results for voice matching.

⁠Directory Structure
speakers/
  your_voice_name/
    model.pth          # Trained So-VITS-SVC model (~50-200MB)
    config.json        # Model configuration
    reference.wav      # Optional: sample of target voice
⁠Adding Trained Voice to Container
  1. Create voice directory: mkdir -p /path/to/speakers/my_voice
  2. Copy trained model files:
   cp /path/to/training/output/model.pth /path/to/speakers/my_voice/
   cp /path/to/training/output/config.json /path/to/speakers/my_voice/
  1. Restart container: docker restart piper-synthesized

Your voice will be available as my_voice in Home Assistant via Wyoming protocol.

⁠Configuration

⁠Environment Variables
VariableDefaultDescription
PIPER_VOICEen_US-lessac-mediumBase Piper voice for synthesis
PIPER_LENGTH_SCALE1.0Speech speed (0.5-2.0)
SILENCE_PAD_MS0Silence padding in milliseconds
SPEAKERS_DIR/app/speakersDirectory containing voice models
⁠Available Piper Voices
  • en_US-lessac-medium - Clear American English (default)
  • en_US-amy-medium - Female American English
  • en_US-ryan-medium - Male American English
  • en_GB-alan-medium - British English

⁠Home Assistant Integration

The container runs a Wyoming protocol server that Home Assistant can discover and use.

⁠Adding to Home Assistant
  1. Settings → Devices & Services → Add Integration
  2. Search for "Wyoming Protocol"
  3. Enter:
    • Host: your-server-ip
    • Port: 10200
  4. Select your custom voice from the dropdown
⁠Using in Automations
action:
  - service: tts.speak
    target:
      entity_id: media_player.living_room
    data:
      message: "Hello! This is my custom cloned voice."
      media_player_entity_id: media_player.living_room
      options:
        voice: my_voice  # Your custom voice name
⁠Using with Voice Assistants

Works with Home Assistant's Assist voice assistant:

  1. Settings → Voice Assistants → Add Assistant
  2. Select Wyoming TTS as speech output
  3. Choose your custom voice

⁠Performance

  • Synthesis time: ~15-20 seconds total per sentence
  • Piper synthesis: ~1-2 seconds (very fast)
  • Voice conversion: ~13-18 seconds (GPU-accelerated)
  • VRAM usage: ~2-4GB during inference
  • Quality: Good voice match with some robotic artifacts from Piper

⁠Troubleshooting

No voices available

  • Check speaker directory mount: docker exec piper-synthesized ls /app/speakers
  • Verify model.pth and config.json exist in each voice folder
  • Check logs: docker logs piper-synthesized

Audio sounds robotic

  • This is a known limitation of Piper synthesis
  • Increase PIPER_LENGTH_SCALE to 1.1 for slightly more natural cadence
  • Add SILENCE_PAD_MS: "200" for better pacing
  • Try different base Piper voice

Slow synthesis

  • Verify GPU access: docker exec piper-synthesized nvidia-smi
  • Check VRAM isn't exhausted (need ~4GB free)
  • Ensure GPU drivers are properly installed on host

Poor voice quality

  • Ensure So-VITS-SVC model was trained with sufficient clean audio (5+ minutes)
  • Try retraining with higher quality recordings
  • Experiment with different base Piper voices

Wyoming connection fails

  • Verify port 10200 is not blocked by firewall
  • Check container logs: docker logs piper-synthesized
  • Ensure container has started successfully

⁠Training Resources

⁠Technical Details

  • Protocol: Wyoming for Home Assistant compatibility
  • TTS Engine: Piper (fast neural synthesis)
  • Voice Conversion: So-VITS-SVC (Singing Voice Conversion)
  • Streaming: Chunked audio delivery
  • Port: 10200 (Wyoming protocol)

⁠Limitations

  • Robotic quality: Piper synthesis has robotic characteristics that So-VITS can't fully remove
  • Speed: Slower than pure Piper due to voice conversion step
  • Training required: Custom voices require training a So-VITS model
  • GPU required: Voice conversion needs GPU for acceptable performance

Tag summary

Content type

Image

Digest

sha256:1ab02ceeb…

Size

4.8 GB

Last updated

7 months ago

docker pull nullableeth/piper-sovits-wyoming