A high-performance Text-to-Speech API using Chatterbox TTS model powered by LitServe.
4.8K
Chatterbox TTS is a state-of-the-art, open-source text-to-speech system developed by Resemble AIโ . It supports zero-shot voice cloning, emotion control, and high-quality audio generation, all MIT-licensed and fully production-ready.
This Docker image wraps Chatterbox TTS in a scalable API using LitServeโ , making it super easy to deploy, clone voices, and generate speech from text.
docker run -p 8000:8000 bhimrazy/chatterbox-tts:v0.1.0
docker run --gpus all -p 8000:8000 bhimrazy/chatterbox-tts:v0.1.0
โ Requires NVIDIA Docker setup for GPU support
POST /speech โ Generate speech from text, with optional voice cloningimport requests
url = "http://127.0.0.1:8000/speech"
data = {
"text": "Hello! This is a test of the Chatterbox TTS API.",
"exaggeration": 0.5,
"cfg": 0.5,
"temperature": 0.8
}
response = requests.post(url, json=data)
with open("output.wav", "wb") as f:
f.write(response.content)
data = {
"text": "This will sound like the reference speaker!",
"audio_prompt": "path/to/reference/audio.wav",
"exaggeration": 0.7,
"cfg": 0.3
}
response = requests.post(url, json=data)
import base64
with open("reference.wav", "rb") as f:
audio_data = f.read()
audio_base64 = base64.b64encode(audio_data).decode('utf-8')
data = {
"text": "Voice cloning with base64 encoded audio!",
"audio_prompt": audio_base64,
"exaggeration": 0.6,
"cfg": 0.4
}
response = requests.post("http://127.0.0.1:8000/speech", json=data)
curl -X POST http://127.0.0.1:8000/speech \
-H "Content-Type: application/json" \
-d '{
"text": "Hello from curl!",
"exaggeration": 0.5,
"cfg": 0.5,
"temperature": 0.8
}' --output output.wav
| Parameter | Type | Range | Description |
|---|---|---|---|
text | string | 1โ500 chars | Input text to synthesize |
audio_prompt | string | optional | File path or base64-encoded reference |
exaggeration | float | 0.0โ1.0 | Emotion intensity |
cfg | float | 0.0โ1.0 | Classifier-free guidance (controls quality/creativity) |
temperature | float | 0.0โ1.0 | Sampling randomness (affects variety) |
If you're working on voice AI, creative tools, or AI-powered agents โ letโs connect! You can follow or fork the repo, and feel free to open issues or contribute enhancements.
Content type
Image
Digest
sha256:d7bf04585โฆ
Size
3.3 GB
Last updated
over 1 year ago
docker pull bhimrazy/chatterbox-tts:v0.1.0