Sign inSign up

gabrielsv01/chatterbox-tts-ptbr

By gabrielsv01

•Updated 14 days ago

Image
Machine learning & AI
0

368

gabrielsv01/chatterbox-tts-ptbr repository overview

A self-hosted REST API around Resemble AI's Chatterbox multilingual text-to-speech model, running the ResembleAI/Chatterbox-Multilingual-pt-br checkpoint (Chatterbox fine-tuned for Brazilian Portuguese). CPU-only, no external API key required.

⚠️ The model uses ~6.5GB of RAM when loaded. By default this app only loads it on the first /tts request and frees it again after 30 minutes idle (the UI/Swagger stay usable the whole time at ~350MB) — see LAZY_LOAD_ENABLED / IDLE_UNLOAD_ENABLED below to change that. First run also downloads ~5GB of weights from Hugging Face.

⚙️ A web UI to try every endpoint lives at "umbrel.local:5160/" — Swagger/ OpenAPI docs are at "umbrel.local:5160/docs".

Endpoints:

  • POST /tts — generate speech from text. Exposed parameters: text, language_id (defaults to "pt", any of the 23 languages the base model supports), exaggeration (emotion intensity), cfg_weight (guidance strength), temperature, repetition_penalty, min_p, top_p, seed (for reproducible output), plus an optional audio_prompt file or a saved voice_name for zero-shot voice cloning. Returns a WAV file.

  • POST /voices and GET/DELETE /voices — save a short reference clip once and reuse it by name in later /tts calls instead of re-uploading it.

  • GET /languages — list of supported language codes.

  • POST /tts/jobs and GET /tts/jobs/{id} — same parameters as /tts, but async with a live progress percentage; used by the web UI's progress bar.

  • GET /outputs, GET/DELETE /outputs/{filename} — history of generated audio, auto-deleted after OUTPUT_RETENTION_DAYS (default 1 day).

  • GET /health — status is "loading" (downloading ~5GB on first ever load, otherwise just reading the cached checkpoint back into RAM), "ready", "idle" (not currently loaded — either never used yet or freed after being idle), or "error".

⚙️ LAZY_LOAD_ENABLED (default true): don't load the model at container startup, only on the first /tts/POST /tts/jobs request. IDLE_UNLOAD_ENABLED (default true) + IDLE_UNLOAD_MINUTES (default 30): free it again after that long with no generation activity. Either request after a cold/idle state gets a 503 while it (re)loads, same as a fresh install. Set either to "false" in the app's environment to keep the model loaded at all times instead.

Quality is best for Portuguese; other languages are inherited from the base multilingual model and may vary.

Tag summary

Content type

Image

Digest

sha256:182a979f2…

Size

746.6 MB

Last updated

14 days ago

docker pull gabrielsv01/chatterbox-tts-ptbr