Sign inSign up

nullableeth/qwentts-wyoming

By nullableeth

•Updated 7 months ago

Zero-shot voice cloning TTS with Qwen3-TTS. Wyoming protocol for Home Assistant. GPU-accelerated.

Image
Machine learning & AI
0

1.9K

nullableeth/qwentts-wyoming repository overview

⁠TLDR: Fourth attempt at voice synthesis/changing/cloning/custom voice, this one is excellent, sounds very close to reference sample the prosody is better than most and sounds close to human speech. This would be my choice all day hands down, the problem it uses a lot of VRAM and is still pretty slow comparatively. I still measure the competitors against this model though.

⁠Qwen3-TTS Wyoming Server

GPU-accelerated Qwen3-TTS server with Wyoming protocol support for Home Assistant and other applications.

⁠Overview

This container provides access to all three Qwen3-TTS model types:

  • CustomVoice: 9 premium pre-trained voices
  • VoiceDesign: Generate voices from text descriptions
  • Base: Clone voices from 3-second audio samples

⁠Features

  • ✅ GPU-accelerated inference with CUDA
  • ✅ Wyoming protocol for Home Assistant integration
  • ✅ Dynamic model and voice selection via environment variables
  • ✅ Supports 11 languages
  • ✅ Multiple voice cloning support (Base models)

⁠Requirements

  • Docker with NVIDIA Container Toolkit
  • NVIDIA GPU with CUDA 12.x support
  • ~4GB GPU VRAM

⁠Quick Start

⁠Option 1: CustomVoice (9 Premium Voices)

Use pre-trained voices with high quality.

services:
  wyoming-qwentts:
    container_name: wyoming-qwentts
    image: nullableeth/qwentts-wyoming:latest
    restart: unless-stopped
    ports:
      - "10204:10800"
    environment:
      QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice"
      QWEN_VOICE: "ryan"  # serena, vivian, uncle_fu, ryan, aiden, ono_anna, sohee, eric, dylan
      QWEN_LANGUAGE: "english"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

Available voices: serena, vivian, uncle_fu, ryan, aiden, ono_anna, sohee, eric, dylan

⁠Option 2: VoiceDesign (Text-to-Voice)

Generate voices from natural language descriptions.

services:
  wyoming-qwentts:
    container_name: wyoming-qwentts
    image: nullableeth/qwentts-wyoming:latest
    restart: unless-stopped
    ports:
      - "10204:10800"
    environment:
      QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign"
      VOICE_DESIGN_DESC: "A female voice with a soft and gentle tone"
      QWEN_LANGUAGE: "english"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

Note: VoiceDesign is only available in the 1.7B model.

⁠Option 3: Base (Voice Cloning from Samples)

Clone voices from 3-second audio reference samples.

services:
  wyoming-qwentts:
    container_name: wyoming-qwentts
    image: nullableeth/qwentts-wyoming:latest
    restart: unless-stopped
    ports:
      - "10204:10800"
    environment:
      QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-Base"
      QWEN_LANGUAGE: "english"
    volumes:
      - ./voices:/app/voices:ro
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

Voice directory structure:

voices/
├── ella/
│   ├── reference.wav  (3+ seconds of clean audio)
│   └── reference.txt  (transcript of the audio)
└── morgan/
    ├── reference.wav
    └── reference.txt

Each directory name becomes a selectable voice in Wyoming (e.g., "ella", "morgan").

⁠Configuration

⁠Environment Variables
VariableDefaultOptionsDescription
QWEN_MODELQwen/Qwen3-TTS-12Hz-0.6B-BaseSee models belowModel to load
QWEN_LANGUAGEenglishSee languages belowSynthesis language
QWEN_VOICEAuto-selectSpeaker nameVoice for CustomVoice models
VOICE_DESIGN_DESCDefault descriptionAny textVoice description for VoiceDesign
⁠Available Models
ModelSizeTypeFeatures
Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice0.6BCustomVoice9 premium voices, faster
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice1.7BCustomVoice9 premium voices, higher quality
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign1.7BVoiceDesignGenerate from description
Qwen/Qwen3-TTS-12Hz-0.6B-Base0.6BBaseVoice cloning, faster
Qwen/Qwen3-TTS-12Hz-1.7B-Base1.7BBaseVoice cloning, higher quality
⁠Supported Languages
  • auto - Auto-detect
  • english
  • chinese
  • japanese
  • korean
  • german
  • french
  • russian
  • portuguese
  • spanish
  • italian
⁠CustomVoice Speakers

Male voices: ryan, aiden, eric, dylan, uncle_fu
Female voices: serena, vivian, ono_anna, sohee

⁠Home Assistant Integration

Add to configuration.yaml:

tts:
  - platform: wyoming
    host: 192.168.1.100  # Your Docker host IP
    port: 10204
    voice: ryan  # For CustomVoice models

Test in Home Assistant:

service: tts.speak
data:
  entity_id: media_player.living_room
  message: "Hello from Qwen3 TTS"
  options:
    voice: serena  # Choose any available voice

⁠Performance

GPU (RTX 4070):

ModelVoice TypeGeneration Time*VRAM Usage
0.6B CustomVoiceryan~8-12s~3.7 GB
1.7B CustomVoiceryan~8-12s~4.0 GB
0.6B Basecloned~8-12s~3.7 GB
1.7B Basecloned~8-12s~4.0 GB

*For a typical 20-word sentence

Note: Qwen models are optimized for quality over speed. For faster generation (<1s), consider using Kokoro TTS with RVC voice conversion.

⁠Command Line Usage

Using Wyoming client tools:

# Install wyoming client
pip install wyoming

# Generate speech
echo "Hello world" | wyoming-client \
  --host localhost \
  --port 10204 \
  --voice ryan \
  --output output.wav

⁠Advanced Usage

⁠Chaining with RVC Voice Conversion

For best results, use Base model for prosody, then apply RVC for custom voice:

services:
  qwen-tts:
    image: nullableeth/qwentts-wyoming:latest
    ports:
      - "10204:10800"
    environment:
      QWEN_MODEL: "Qwen/Qwen3-TTS-12Hz-1.7B-Base"
    volumes:
      - ./voices:/app/voices:ro
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
  
  rvc-converter:
    image: nullableeth/rvc-wyoming:latest
    ports:
      - "10900:10900"
    environment:
      TTS_HOST: "qwen-tts"
      TTS_PORT: "10800"
      TTS_VOICE: "ella"  # Your cloned voice
      MODEL_NAME: "your_model.pth"
    volumes:
      - ./rvc_models:/models:ro
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

Point Home Assistant to port 10900 for RVC-converted output.

⁠Troubleshooting

⁠Container exits with "Base model requires voice cloning"

Cause: Using a Base model without mounting /app/voices.

Solution: Either:

  1. Mount voice samples to /app/voices, OR
  2. Switch to CustomVoice or VoiceDesign model
⁠"No speakers available" error

Cause: CustomVoice model failed to load speaker list.

Solution: Check container logs:

docker logs wyoming-qwentts | grep -i speaker
⁠GPU not being used

Cause: NVIDIA Container Toolkit not installed or GPU not accessible.

Solution:

  1. Verify GPU access:
   docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi
  1. Check container logs for "Using GPU" message
⁠Slow generation (>30 seconds)

Cause: Running on CPU instead of GPU.

Solution: Ensure GPU is accessible and check logs for "Using GPU with flash-attention" message.

⁠Voice has unexpected accent

Cause: CustomVoice speakers may have non-native accents.

Solution: Try different speakers (ryan, aiden, serena, vivian) or use Base model with native English reference audio.

⁠Technical Details

⁠Architecture
  • Base Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
  • Model: Qwen3-TTS (0.6B or 1.7B parameters)
  • Attention: Flash Attention 2 for GPU acceleration
  • Protocol: Wyoming for Home Assistant compatibility
  • Audio Output: 12kHz sample rate, 16-bit PCM
⁠Voice Cloning (Base Models)

Base models use a 3-second reference audio sample to clone voice characteristics:

  • x_vector_only_mode=False: Uses both speaker embedding and in-context learning (ICL)
  • Requires both reference.wav (3+ seconds) and reference.txt (transcript)
  • Each subdirectory in /app/voices/ becomes a selectable voice
⁠Model Downloads

Models are downloaded on first run from HuggingFace:

  • 0.6B models: ~2-3 GB download
  • 1.7B models: ~4-5 GB download

Models are cached in the container at /root/.cache/huggingface/.

⁠Building From Source

git clone <your-repo>
cd qwentts-wyoming/build
docker build -t qwentts-wyoming:latest .

⁠License

This container packages:

  • Qwen3-TTS (Model license: see Qwen GitHub⁠)
  • Wyoming Protocol (MIT)
  • PyTorch (BSD-style)

⁠Credits

⁠Support

For issues specific to this container, please open an issue on GitHub. For Qwen3-TTS model issues, see the official Qwen repository⁠.

⁠Comparison with Other TTS

ModelSpeedQualityVoice OptionsBest For
Kokoro⚡ Very Fast (0.2-0.6s)High54 voicesReal-time applications
Qwen3-TTSSlow (8-12s)Very High9 voices + cloningQuality over speed
PiperFast (1-2s)Medium100+ voicesResource-constrained
StyleTTS2Slow (10-20s)Very HighVoice cloningStudio quality

Recommendation: For sub-second generation with custom voices, use Kokoro + RVC instead of Qwen3-TTS.

Tag summary

Content type

Image

Digest

sha256:81477b29c…

Size

7.4 GB

Last updated

7 months ago

docker pull nullableeth/qwentts-wyoming