Sign inSign up

bhimrazy/chatterbox-tts

By bhimrazy

โ€ขUpdated over 1 year ago

A high-performance Text-to-Speech API using Chatterbox TTS model powered by LitServe.

Image
Machine learning & AI
Web servers
2

4.8K

bhimrazy/chatterbox-tts repository overview

โ ๐Ÿง  Production-Ready Text-to-Speech API with Chatterbox TTS & LitServe

Chatterbox TTS is a state-of-the-art, open-source text-to-speech system developed by Resemble AIโ . It supports zero-shot voice cloning, emotion control, and high-quality audio generation, all MIT-licensed and fully production-ready.

This Docker image wraps Chatterbox TTS in a scalable API using LitServeโ , making it super easy to deploy, clone voices, and generate speech from text.


โ ๐Ÿ“ฆ Quick Start with Docker

โ Pull & Run the Container
docker run -p 8000:8000 bhimrazy/chatterbox-tts:v0.1.0
โ Enable GPU Acceleration
docker run --gpus all -p 8000:8000 bhimrazy/chatterbox-tts:v0.1.0

โœ… Requires NVIDIA Docker setup for GPU support


โ ๐ŸŽฏ API Endpoints

  • POST /speech โ€“ Generate speech from text, with optional voice cloning

โ ๐Ÿงช Example Usage

โ ๐Ÿงต Python: Basic Text-to-Speech
import requests

url = "http://127.0.0.1:8000/speech"
data = {
    "text": "Hello! This is a test of the Chatterbox TTS API.",
    "exaggeration": 0.5,
    "cfg": 0.5,
    "temperature": 0.8
}

response = requests.post(url, json=data)

with open("output.wav", "wb") as f:
    f.write(response.content)

โ ๐Ÿงฌ Voice Cloning with Local Audio File
data = {
    "text": "This will sound like the reference speaker!",
    "audio_prompt": "path/to/reference/audio.wav",
    "exaggeration": 0.7,
    "cfg": 0.3
}

response = requests.post(url, json=data)

โ ๐Ÿ“ฆ Voice Cloning with Base64 Audio
import base64

with open("reference.wav", "rb") as f:
    audio_data = f.read()
audio_base64 = base64.b64encode(audio_data).decode('utf-8')

data = {
    "text": "Voice cloning with base64 encoded audio!",
    "audio_prompt": audio_base64,
    "exaggeration": 0.6,
    "cfg": 0.4
}

response = requests.post("http://127.0.0.1:8000/speech", json=data)

โ ๐Ÿงพ Curl Example
curl -X POST http://127.0.0.1:8000/speech \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Hello from curl!",
    "exaggeration": 0.5,
    "cfg": 0.5,
    "temperature": 0.8
  }' --output output.wav

โ ๐Ÿ”— Resources


โ ๐Ÿ”ง Parameters Overview

ParameterTypeRangeDescription
textstring1โ€“500 charsInput text to synthesize
audio_promptstringoptionalFile path or base64-encoded reference
exaggerationfloat0.0โ€“1.0Emotion intensity
cfgfloat0.0โ€“1.0Classifier-free guidance (controls quality/creativity)
temperaturefloat0.0โ€“1.0Sampling randomness (affects variety)

โ ๐Ÿ› ๏ธ Use Cases

  • ๐ŸŽฎ Game characters with emotional dialogue
  • ๐Ÿง‘โ€๐Ÿ’ผ AI assistants that sound real
  • ๐ŸŽ™๏ธ Podcast / audiobook narration
  • ๐ŸŽฌ Automated video voice-overs
  • ๐ŸŒ Multilingual content with consistent branding

โ ๐Ÿ‘‹ Get Involved

If you're working on voice AI, creative tools, or AI-powered agents โ€” letโ€™s connect! You can follow or fork the repo, and feel free to open issues or contribute enhancements.

Tag summary

Content type

Image

Digest

sha256:d7bf04585โ€ฆ

Size

3.3 GB

Last updated

over 1 year ago

docker pull bhimrazy/chatterbox-tts:v0.1.0