VoxCPMv2 WebUI & API – Docker TTS with offline-ready voices and GPU support. (30 languages)
1.5K
Easy-to-run VoxCPM2 text-to-speech Docker images with a browser UI and HTTP API included.
This Hangry Labs fork is built for people who want realistic multilingual text to speech, voice design, and voice cloning without a long setup. Install Docker, run one command, open the local UI, or call the API from your own application.
VoxCPM2 supports highly realistic voice cloning. Do not use this image for unauthorized voice cloning, impersonation, fraud, harassment, scams, or any illegal or unethical activity. Only clone voices when you have the rights and consent to do so.
Run with NVIDIA GPU support:
docker run -p 8808:8808 --gpus all hangrylabs/voxcpmtts:v0.1
Run on CPU:
docker run -p 8808:8808 -e VOXCPM_DEVICE=cpu hangrylabs/voxcpmtts:v0.1
Run on a specific GPU:
docker run -p 8808:8808 --gpus "device=1" -e CUDA_VISIBLE_DEVICES=1 hangrylabs/voxcpmtts:v0.1
Then open:
The standard vX.Y image is the full baked image with VoxCPM2 model assets plus the denoiser and ASR support assets included for offline-friendly use after the image is pulled.
The runtime defaults to VOXCPM_OPTIMIZE=0 so the slim image does not need a C compiler for first-run Triton compilation. Set VOXCPM_OPTIMIZE=1 only when you want to test compiled inference.
Tiny tags use the vX.Y_tiny pattern. They keep runtime dependencies but skip baked Hugging Face model assets, and are intended for persistent-volume workflows where the cache is warmed on first online use:
docker run -p 8808:8808 --gpus all -v voxcpmtts_hf_cache:/app/.cache/huggingface hangrylabs/voxcpmtts:v0.1_tiny
voice, use_gpu, /tts/voices, /tts/speakers, /tts/stream-formats, and /tts/streamDefault API behavior returns WAV:
curl -X POST "http://localhost:8808/tts/generate" \
-H "Content-Type: application/json" \
-d '{"text":"Hello from Hangry Labs VoxCPMTTS","language":"English"}' \
-o hello.wav
Request MP3 when you want compact output:
curl -X POST "http://localhost:8808/tts/generate" \
-H "Content-Type: application/json" \
-d '{"text":"Hello from Hangry Labs VoxCPMTTS","language":"English","output_format":"mp3"}' \
-o hello.mp3
Voice design:
curl -X POST "http://localhost:8808/tts/generate" \
-H "Content-Type: application/json" \
-d '{"text":"This is a custom designed voice.","control":"young female, warm, gentle, slightly smiling","output_format":"mp3"}' \
-o designed.mp3
Voice cloning can be called with a reference audio path that is visible inside the container:
curl -X POST "http://localhost:8808/tts/generate" \
-H "Content-Type: application/json" \
-d '{"text":"This voice follows the reference sample.","ref_audio":"/data/ref.wav","output_format":"mp3"}' \
-o cloned.mp3
Transcript-guided cloning:
curl -X POST "http://localhost:8808/tts/generate" \
-H "Content-Type: application/json" \
-d '{"text":"The model continues from the reference voice.","ref_audio":"/data/ref.wav","ref_text":"Transcript of the reference audio.","output_format":"mp3"}' \
-o ultimate.mp3
Health check:
curl http://localhost:8808/tts/ping
API docs are available at:
http://localhost:8808/tts/docs
v0.1vX.YvX.Y_tinyThis is an independently maintained Hangry Labs packaging and serving fork of the original VoxCPM project by OpenBMB, ModelBest, THUHCSI, and contributors:
https://github.com/OpenBMB/VoxCPM
License and attribution are preserved in the repository. Original VoxCPM copyright remains with the upstream authors; Hangry Labs maintains the Docker packaging, Web UI/API integration, documentation, release tooling, and related modifications in this fork.
Content type
Image
Digest
sha256:bbf3a0e45…
Size
9 GB
Last updated
5 months ago
docker pull hangrylabs/voxcpmtts