Sign inSign up

nullableeth/rvc-trainer

By nullableeth

•Updated 7 months ago

Pre-patched RVC training for voice conversion models. PyTorch 2.6+ compatible. GPU-accelerated.

Image
Machine learning & AI
0

1.4K

nullableeth/rvc-trainer repository overview

⁠TLDR: Fifth attempt at voice synthesis/changing/cloning/custom voice, this is just a training container for training RVC models on reference audio. No reference.txt needed for this one but setting up the training environment was a pain so I posted the container. You have to train a model on 5-10min data and take 9+ hours.

⁠RVC Trainer - Voice Conversion Model Training

Pre-patched RVC (Retrieval-based Voice Conversion) training environment with all compatibility fixes for PyTorch 2.6+ and modern dependencies.

⁠What This Is

Training container for RVC voice conversion models. Use this to train custom voice models from audio samples, then deploy them with a separate inference container for voice conversion or Text-to-Speech pipelines.

⁠Features

  • ✅ Pre-patched - All PyTorch 2.6+ and matplotlib compatibility fixes applied
  • 🔧 Ready to train - No manual patching required
  • 🎯 RVC v2 - Latest Retrieval-based Voice Conversion architecture
  • 📦 Complete environment - All dependencies included
  • 🚀 GPU accelerated - NVIDIA CUDA support

⁠Use Cases

Train RVC models for:

  • Voice conversion - Convert any voice to sound like your target voice
  • TTS pipelines - Use with Piper, Kokoro, or other TTS engines for custom voices
  • Real-time voice changing - Deploy trained models for live voice conversion
  • Accent preservation - Maintain regional accents better than pure TTS

⁠Quick Start

⁠Training a Voice Model
# Prepare your audio (10-15 minutes of clean speech recommended)
mkdir -p training/input
cp your_voice.wav training/input/

# Start training container
docker run -it --rm --gpus all \
  --shm-size=8g \
  -v $(pwd)/training:/app/training \
  nullableeth/rvc-trainer:latest \
  bash

# Inside container - preprocess audio
cd /app/rvc
python3 infer/modules/train/preprocess.py \
  /app/training/input \
  40000 \
  4 \
  /app/training/my_voice \
  False \
  1.0

# Extract features (run all 4 parts)
for part in 0 1 2 3; do
  python3 infer/modules/train/extract_feature_print.py \
    cuda:0 4 $part /app/training/my_voice v2 False
done

# Extract pitch
python3 infer/modules/train/extract/extract_f0_print.py \
  /app/training/my_voice \
  4 \
  harvest

# Generate filelist
python3 << 'PYEOF'
import os
exp_dir = '/app/training/my_voice'
feature_files = sorted([f.replace('.npy', '') for f in os.listdir(f'{exp_dir}/3_feature768') if f.endswith('.npy')])
with open(f'{exp_dir}/filelist.txt', 'w') as f:
    for basename in feature_files:
        gt_wav = f"{exp_dir}/0_gt_wavs/{basename}.wav"
        phone = f"{exp_dir}/3_feature768/{basename}.npy"
        f0 = f"{exp_dir}/2a_f0/{basename}.wav.npy"
        f0nsf = f"{exp_dir}/2b-f0nsf/{basename}.wav.npy"
        if all(os.path.exists(p) for p in [gt_wav, phone, f0, f0nsf]):
            f.write(f"{gt_wav}|{phone}|{f0}|{f0nsf}|0\n")
PYEOF

# Create config (see full example below)
# Then train...
python3 infer/modules/train/train.py \
  -e /app/training/my_voice \
  -sr 40k \
  -f0 1 \
  -bs 2 \
  -g 0 \
  -te 200 \
  -se 50 \
  -pg assets/pretrained_v2/f0G40k.pth \
  -pd assets/pretrained_v2/f0D40k.pth \
  -l 0 \
  -c 0 \
  -sw 0 \
  -v v2
⁠Config File Template

Create training/my_voice/config.json:

{
  "train": {
    "log_interval": 200,
    "seed": 1234,
    "epochs": 200,
    "learning_rate": 1e-4,
    "betas": [0.8, 0.99],
    "eps": 1e-9,
    "batch_size": 2,
    "fp16_run": false,
    "lr_decay": 0.999875,
    "segment_size": 12800,
    "init_lr_ratio": 1,
    "warmup_epochs": 0,
    "c_mel": 45,
    "c_kl": 1.0
  },
  "data": {
    "max_wav_value": 32768.0,
    "sampling_rate": 40000,
    "filter_length": 2048,
    "hop_length": 400,
    "win_length": 2048,
    "n_mel_channels": 125,
    "mel_fmin": 0.0,
    "mel_fmax": null
  },
  "model": {
    "inter_channels": 192,
    "hidden_channels": 192,
    "filter_channels": 768,
    "n_heads": 2,
    "n_layers": 6,
    "kernel_size": 3,
    "p_dropout": 0,
    "resblock": "1",
    "resblock_kernel_sizes": [3,7,11],
    "resblock_dilation_sizes": [[1,3,5], [1,3,5], [1,3,5]],
    "upsample_rates": [10,10,2,2],
    "upsample_initial_channel": 512,
    "upsample_kernel_sizes": [16,16,4,4],
    "use_spectral_norm": false,
    "gin_channels": 256,
    "spk_embed_dim": 109
  },
  "version": "v2",
  "sr": "40k",
  "if_f0": 1,
  "spk": {
    "my_voice": 0
  }
}

⁠Requirements

  • Audio: 10-15 minutes of clean speech (WAV format recommended)
  • GPU: NVIDIA GPU with 12GB+ VRAM
  • Disk: ~5GB for models and training data
  • Time: 8-12 hours for 200 epochs

⁠Training Output

Trained models saved as checkpoints:

  • training/my_voice/G_50.pth (epoch 50)
  • training/my_voice/G_100.pth (epoch 100)
  • training/my_voice/G_150.pth (epoch 150)
  • training/my_voice/G_200.pth (epoch 200)

Use these .pth files in your inference container for voice conversion.

⁠Resuming Training

To continue training from a checkpoint:

# Change -l 0 to -l 1 and increase -te
python3 infer/modules/train/train.py \
  -e /app/training/my_voice \
  -sr 40k \
  -f0 1 \
  -bs 2 \
  -g 0 \
  -te 500 \
  -se 50 \
  -pg assets/pretrained_v2/f0G40k.pth \
  -pd assets/pretrained_v2/f0D40k.pth \
  -l 1 \
  -c 0 \
  -sw 0 \
  -v v2

⁠What's Patched

This image includes fixes for:

  • PyTorch 2.6+ weights_only parameter in fairseq
  • Matplotlib tostring_rgb() deprecation
  • Fairseq checkpoint loading compatibility

⁠Not Included

This is a training-only container. For inference/deployment:

  • Build your own inference container with Wyoming protocol
  • Use with TTS engines like Piper, Kokoro, XTTS
  • Integrate into voice conversion pipelines

⁠Training Tips

  • Audio quality: Clean recordings without background noise work best
  • Duration: 10-15 minutes ideal, 5 minutes minimum works but lower quality
  • Preprocessing threshold: Use 1.0 for less aggressive silence removal, 3.0 for more aggressive
  • Epochs: Test checkpoints at 50, 100, 150 - quality plateaus around 100-200
  • Batch size: Reduce to 1 if GPU runs out of memory
  • Checkpoints: Save every 20-50 epochs to test incremental improvements

⁠Common Issues

GPU Out of Memory: Reduce batch size (-bs 1)

Training too slow: Reduce total epochs (-te 100) and test early checkpoints

Poor quality output: Ensure reference audio is clean, increase training data to 15+ minutes

Container disconnects: Use screen or tmux to keep training running during SSH disconnects

Tag summary

Content type

Image

Digest

sha256:e0de4dc80…

Size

5.5 GB

Last updated

7 months ago

docker pull nullableeth/rvc-trainer