Sign inSign up

nullableeth/rvc-wyoming

By nullableeth

•Updated 7 months ago

Modular RVC voice converter for Wyoming TTS servers. Apply custom voices to any TTS. GPU-accelerated

Image
Machine learning & AI
0

1.5K

nullableeth/rvc-wyoming repository overview

⁠RVC Voice Conversion Wyoming Server

Modular RVC (Retrieval-based Voice Conversion) server that wraps any Wyoming TTS service to apply custom voice conversion. Works as a drop-in replacement for the upstream TTS in Home Assistant.

⁠What is This?

This container sits between your client (Home Assistant, etc.) and any Wyoming TTS server (Kokoro, Piper, StyleTTS2, etc.) to apply real-time voice conversion using your trained RVC model. It receives TTS requests, forwards them to an upstream TTS, then converts the voice to match your custom RVC model.

[Home Assistant] → [RVC Container] → [Upstream TTS (Kokoro/Piper/etc.)]
                         ↓
                  [Converted Voice Output]

⁠Features

  • ✅ Works with any Wyoming TTS server (Kokoro, Piper, StyleTTS2, XTTS, etc.)
  • ✅ GPU-accelerated voice conversion
  • ✅ Wyoming protocol compatible
  • ✅ Real-time conversion (~0.2-0.5s overhead)
  • ✅ Preserves prosody and naturalness from source TTS

⁠Requirements

  • Docker with NVIDIA Container Toolkit
  • NVIDIA GPU with CUDA support
  • ~2GB GPU VRAM
  • Trained RVC v2 model (.pth file and config.json)
  • Upstream Wyoming TTS server (e.g., Kokoro, Piper)

⁠Quick Start

⁠1. Prepare Your RVC Model

Place your trained RVC model files in a directory:

./rvc_models/
├── your_model.pth      # Your trained RVC checkpoint
└── config.json         # Training configuration file

Don't have a model? See Training Your Own RVC Model⁠ below.

⁠2. Start Upstream TTS

First, start your source TTS server (example using Kokoro):

docker run -d \
  --name kokoro-tts \
  --gpus all \
  -p 10210:10210 \
  nullableeth/kokoro-wyoming:latest
⁠3. Start RVC Converter
docker run -d \
  --name rvc-converter \
  --gpus all \
  -p 10900:10900 \
  -v /path/to/rvc_models:/models:ro \
  -e TTS_HOST=172.17.0.1 \
  -e TTS_PORT=10210 \
  -e TTS_VOICE=af_sky \
  -e MODEL_NAME=your_model.pth \
  nullableeth/rvc-wyoming:latest

Note: Replace 172.17.0.1 with your Docker host IP or container name if using docker-compose.

services:
  kokoro-tts:
    container_name: kokoro-tts
    image: nullableeth/kokoro-wyoming:latest
    restart: unless-stopped
    ports:
      - "10210:10210"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

  rvc-converter:
    container_name: rvc-converter
    image: nullableeth/rvc-wyoming:latest
    restart: unless-stopped
    ports:
      - "10900:10900"
    volumes:
      - ./rvc_models:/models:ro
    environment:
      WYOMING_PORT: "10900"
      MODEL_NAME: "your_model.pth"
      TTS_HOST: "kokoro-tts"  # Container name
      TTS_PORT: "10210"
      TTS_VOICE: "af_sky"
    depends_on:
      - kokoro-tts
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

⁠Environment Variables

VariableRequiredDefaultDescription
WYOMING_PORTNo10900Port for RVC server to listen on
MODEL_NAMEYesG_81400.pthFilename of your RVC model in /models
TTS_HOSTYeslocalhostHostname/IP of upstream Wyoming TTS server
TTS_PORTYes10210Port of upstream Wyoming TTS server
TTS_VOICENoaf_skyVoice to request from upstream TTS

⁠Home Assistant Integration

Point Home Assistant to the RVC port instead of the direct TTS port:

tts:
  - platform: wyoming
    host: 192.168.1.100
    port: 10900  # RVC converter port, NOT the TTS port

Home Assistant will send requests to RVC, which automatically forwards to the upstream TTS and converts the voice.

⁠Performance

Typical pipeline latency (Kokoro + RVC on RTX 4070):

ComponentTimeNotes
Upstream TTS (Kokoro GPU)~0.2-0.6sDepends on TTS engine
RVC Conversion~0.2-0.5sFirst request ~3s (model loading)
Total~0.4-1.1sReal-time capable

⁠Training Your Own RVC Model

This container requires a pre-trained RVC v2 model. Use the companion training container to train your model:

⁠Easy Training with Docker
docker run -it --rm \
  --gpus all \
  -v $(pwd)/dataset:/app/dataset \
  -v $(pwd)/logs:/app/logs \
  nullableeth/rvc-trainer:latest

This launches an interactive training container. See the RVC Trainer documentation⁠ for detailed instructions.

⁠Quick Training Steps
  1. Prepare audio samples (5-15 minutes of clean voice audio recommended)
  2. Run the trainer container (pulls nullableeth/rvc-trainer)
  3. Extract features and train (follow interactive prompts)
  4. Export model - Copy the .pth file and config.json from /app/logs/your_model/
  5. Use with this container - Mount the exported files to /models

Alternative: Use the RVC-WebUI⁠ directly for a web interface.

⁠Compatible Upstream TTS Servers

This container works with any Wyoming protocol TTS server:

  • ✅ Kokoro⁠ (Recommended - fast, natural, 54 voices)
  • ✅ Piper⁠ (Fast, many voices)
  • ✅ StyleTTS2 (High quality, slower)
  • ✅ XTTS (Voice cloning capable)
  • ✅ Any Wyoming-compatible TTS

⁠Troubleshooting

⁠No audio output or errors
  1. Check upstream TTS is running:
   curl http://TTS_HOST:TTS_PORT
  1. Verify GPU access:
   docker exec rvc-converter python3 -c "import torch; print('CUDA:', torch.cuda.is_available())"

Should output: CUDA: True

  1. Check model files exist:
   docker exec rvc-converter ls -la /models

Should show your .pth file and config.json

⁠Slow first request (~5 seconds)

This is normal - RVC loads the model and RMVPE on first use. Subsequent requests are fast (~0.2-0.5s).

⁠Voice sounds distorted or robotic
  • Ensure your RVC model was trained with sufficient clean audio (5-15 minutes minimum)
  • Try adjusting index_rate or protect parameters in the wrapper code
  • Verify the upstream TTS sample rate matches your RVC training sample rate
  • Retrain with higher quality source audio
⁠"Connection refused" to upstream TTS
  • If using docker-compose, use container name as TTS_HOST
  • If using separate containers, ensure network connectivity
  • Check firewall rules on host machine

⁠Advanced Configuration

⁠Custom RVC Parameters

The container uses sensible defaults for RVC conversion. To modify parameters, you'll need to adjust the rvc_convert function in the wrapper:

def rvc_convert(input_path, output_path):
    info, (tgt_sr, audio_opt) = vc.vc_single(
        sid=0,
        input_audio_path=input_path,
        f0_up_key=0,           # Pitch shift (semitones)
        f0_method='rmvpe',     # Pitch extraction method
        index_rate=0.75,       # Feature retrieval strength
        filter_radius=3,       # Median filter for pitch
        resample_sr=0,         # Output sample rate (0=auto)
        rms_mix_rate=0.25,     # Volume envelope mix
        protect=0.33,          # Protect voiceless consonants
    )
⁠Multiple RVC Models

To run multiple RVC converters with different voices:

services:
  rvc-voice1:
    image: nullableeth/rvc-wyoming:latest
    ports:
      - "10901:10900"
    environment:
      MODEL_NAME: "voice1.pth"
      TTS_HOST: "kokoro-tts"
      TTS_VOICE: "af_bella"
    volumes:
      - ./models/voice1:/models:ro
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

  rvc-voice2:
    image: nullableeth/rvc-wyoming:latest
    ports:
      - "10902:10900"
    environment:
      MODEL_NAME: "voice2.pth"
      TTS_HOST: "kokoro-tts"
      TTS_VOICE: "am_adam"
    volumes:
      - ./models/voice2:/models:ro
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

⁠Building From Source

git clone <your-repo>
cd rvc-wyoming
docker build -t rvc-wyoming:latest .

⁠Technical Details

⁠Architecture
  • Base Image: pytorch/pytorch:2.3.1-cuda12.1-cudnn8-runtime
  • RVC Version: v2
  • Pitch Extraction: RMVPE (GPU accelerated)
  • Protocol: Wyoming for TTS communication
⁠Model Requirements

Your RVC model must be:

  • RVC v2 format (.pth checkpoint)
  • Include accompanying config.json from training
  • Trained with consistent sample rate (typically 40kHz or 48kHz)
⁠File Structure
/models/              # Mounted volume (read-only)
├── your_model.pth   # Your trained model
└── config.json      # Training configuration

/app/rvc/            # RVC-WebUI code
├── assets/          # Hubert & RMVPE models (auto-downloaded)
└── ...

/tmp/rvc_models/     # Runtime working directory

⁠License

This container packages:

  • RVC-WebUI (MIT License)
  • Wyoming Protocol (MIT License)
  • PyTorch (BSD License)

⁠Credits

⁠Support

For issues specific to this container, please open an issue on GitHub. For RVC model training questions, see the RVC Trainer documentation⁠.

Tag summary

Content type

Image

Digest

sha256:549b7b938…

Size

8.7 GB

Last updated

7 months ago

docker pull nullableeth/rvc-wyoming