Modular RVC voice converter for Wyoming TTS servers. Apply custom voices to any TTS. GPU-accelerated
1.5K
Modular RVC (Retrieval-based Voice Conversion) server that wraps any Wyoming TTS service to apply custom voice conversion. Works as a drop-in replacement for the upstream TTS in Home Assistant.
This container sits between your client (Home Assistant, etc.) and any Wyoming TTS server (Kokoro, Piper, StyleTTS2, etc.) to apply real-time voice conversion using your trained RVC model. It receives TTS requests, forwards them to an upstream TTS, then converts the voice to match your custom RVC model.
[Home Assistant] → [RVC Container] → [Upstream TTS (Kokoro/Piper/etc.)]
↓
[Converted Voice Output]
.pth file and config.json)Place your trained RVC model files in a directory:
./rvc_models/
├── your_model.pth # Your trained RVC checkpoint
└── config.json # Training configuration file
Don't have a model? See Training Your Own RVC Model below.
First, start your source TTS server (example using Kokoro):
docker run -d \
--name kokoro-tts \
--gpus all \
-p 10210:10210 \
nullableeth/kokoro-wyoming:latest
docker run -d \
--name rvc-converter \
--gpus all \
-p 10900:10900 \
-v /path/to/rvc_models:/models:ro \
-e TTS_HOST=172.17.0.1 \
-e TTS_PORT=10210 \
-e TTS_VOICE=af_sky \
-e MODEL_NAME=your_model.pth \
nullableeth/rvc-wyoming:latest
Note: Replace 172.17.0.1 with your Docker host IP or container name if using docker-compose.
services:
kokoro-tts:
container_name: kokoro-tts
image: nullableeth/kokoro-wyoming:latest
restart: unless-stopped
ports:
- "10210:10210"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
rvc-converter:
container_name: rvc-converter
image: nullableeth/rvc-wyoming:latest
restart: unless-stopped
ports:
- "10900:10900"
volumes:
- ./rvc_models:/models:ro
environment:
WYOMING_PORT: "10900"
MODEL_NAME: "your_model.pth"
TTS_HOST: "kokoro-tts" # Container name
TTS_PORT: "10210"
TTS_VOICE: "af_sky"
depends_on:
- kokoro-tts
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
| Variable | Required | Default | Description |
|---|---|---|---|
WYOMING_PORT | No | 10900 | Port for RVC server to listen on |
MODEL_NAME | Yes | G_81400.pth | Filename of your RVC model in /models |
TTS_HOST | Yes | localhost | Hostname/IP of upstream Wyoming TTS server |
TTS_PORT | Yes | 10210 | Port of upstream Wyoming TTS server |
TTS_VOICE | No | af_sky | Voice to request from upstream TTS |
Point Home Assistant to the RVC port instead of the direct TTS port:
tts:
- platform: wyoming
host: 192.168.1.100
port: 10900 # RVC converter port, NOT the TTS port
Home Assistant will send requests to RVC, which automatically forwards to the upstream TTS and converts the voice.
Typical pipeline latency (Kokoro + RVC on RTX 4070):
| Component | Time | Notes |
|---|---|---|
| Upstream TTS (Kokoro GPU) | ~0.2-0.6s | Depends on TTS engine |
| RVC Conversion | ~0.2-0.5s | First request ~3s (model loading) |
| Total | ~0.4-1.1s | Real-time capable |
This container requires a pre-trained RVC v2 model. Use the companion training container to train your model:
docker run -it --rm \
--gpus all \
-v $(pwd)/dataset:/app/dataset \
-v $(pwd)/logs:/app/logs \
nullableeth/rvc-trainer:latest
This launches an interactive training container. See the RVC Trainer documentation for detailed instructions.
.pth file and config.json from /app/logs/your_model//modelsAlternative: Use the RVC-WebUI directly for a web interface.
This container works with any Wyoming protocol TTS server:
curl http://TTS_HOST:TTS_PORT
docker exec rvc-converter python3 -c "import torch; print('CUDA:', torch.cuda.is_available())"
Should output: CUDA: True
docker exec rvc-converter ls -la /models
Should show your .pth file and config.json
This is normal - RVC loads the model and RMVPE on first use. Subsequent requests are fast (~0.2-0.5s).
index_rate or protect parameters in the wrapper codeTTS_HOSTThe container uses sensible defaults for RVC conversion. To modify parameters, you'll need to adjust the rvc_convert function in the wrapper:
def rvc_convert(input_path, output_path):
info, (tgt_sr, audio_opt) = vc.vc_single(
sid=0,
input_audio_path=input_path,
f0_up_key=0, # Pitch shift (semitones)
f0_method='rmvpe', # Pitch extraction method
index_rate=0.75, # Feature retrieval strength
filter_radius=3, # Median filter for pitch
resample_sr=0, # Output sample rate (0=auto)
rms_mix_rate=0.25, # Volume envelope mix
protect=0.33, # Protect voiceless consonants
)
To run multiple RVC converters with different voices:
services:
rvc-voice1:
image: nullableeth/rvc-wyoming:latest
ports:
- "10901:10900"
environment:
MODEL_NAME: "voice1.pth"
TTS_HOST: "kokoro-tts"
TTS_VOICE: "af_bella"
volumes:
- ./models/voice1:/models:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
rvc-voice2:
image: nullableeth/rvc-wyoming:latest
ports:
- "10902:10900"
environment:
MODEL_NAME: "voice2.pth"
TTS_HOST: "kokoro-tts"
TTS_VOICE: "am_adam"
volumes:
- ./models/voice2:/models:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
git clone <your-repo>
cd rvc-wyoming
docker build -t rvc-wyoming:latest .
pytorch/pytorch:2.3.1-cuda12.1-cudnn8-runtimeYour RVC model must be:
.pth checkpoint)config.json from training/models/ # Mounted volume (read-only)
├── your_model.pth # Your trained model
└── config.json # Training configuration
/app/rvc/ # RVC-WebUI code
├── assets/ # Hubert & RMVPE models (auto-downloaded)
└── ...
/tmp/rvc_models/ # Runtime working directory
This container packages:
For issues specific to this container, please open an issue on GitHub. For RVC model training questions, see the RVC Trainer documentation.
Content type
Image
Digest
sha256:549b7b938…
Size
8.7 GB
Last updated
7 months ago
docker pull nullableeth/rvc-wyoming