Sign inSign up

ffaerber/kokoro-cuda

By ffaerber

•Updated 5 months ago

kokoro-cuda TTS

Image
0

923

ffaerber/kokoro-cuda repository overview

⁠Kokoro CUDA

Docker Image

Minimal Kokoro TTS API with CUDA support for NVIDIA GPUs (Ada Lovelace and newer).

Source: https://github.com/ffaerber/kokoro-cuda⁠

⁠Features

  • Single FastAPI endpoint for text-to-speech
  • Streaming audio (WAV/PCM) and encoded formats (MP3, Opus, FLAC)
  • 49 built-in voices (American/British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, Chinese)
  • Configurable voice, speed, and bitrate
  • Web UI at / — type text, pick a voice, and hear audio instantly
  • Swagger UI at /docs
  • Automatic model + voice download on first start (persisted via Docker volumes)

⁠Quick Start

docker build -t kokoro-cuda .
docker run --gpus all -p 8880:8880 -v kokoro-models:/app/models -v kokoro-voices:/app/voices kokoro-cuda

⁠API

⁠POST /v1/audio/speech
{
  "input": "Hello, this is a test.",
  "voice": "af_heart",
  "speed": 1.0,
  "response_format": "wav",
  "bitrate": "192k"
}
ParameterDefaultOptions
input(required)Any text
voiceaf_heart49 voices — see GET /v1/voices
speed1.00.5 - 2.0
response_formatwavwav, mp3, opus, flac, pcm
bitrate192k128k, 192k, 320k
⁠GET /v1/voices

Returns available voice packs.

⁠GET /health

Returns model status.

⁠Examples

# WAV
curl -X POST http://localhost:8880/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"input": "Hello world"}' -o hello.wav

# MP3 at 320k
curl -X POST http://localhost:8880/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"input": "Hello world", "response_format": "mp3", "bitrate": "320k"}' -o hello.mp3

⁠Web UI

Open http://localhost:8880/ in your browser. Select a voice, adjust speed, type text, and click Speak to hear the audio immediately.

⁠Voices

49 voice packs are downloaded automatically on first start. Naming convention: {lang}{gender}_{name}

PrefixLanguageVoices
af_ / am_American English20
bf_ / bm_British English8
ef_ / em_Spanish3
ff_French1
hf_ / hm_Hindi4
if_ / im_Italian2
jf_ / jm_Japanese5
pf_ / pm_Portuguese3
zf_Chinese4

⁠Benchmark

Run the full benchmark suite (Kokoro + Whisper validation):

cd benchmark
docker compose up --build --abort-on-container-exit --exit-code-from benchmark

Results are saved to benchmark/output/<GPU_NAME>/report.md. See latest results⁠.

⁠Architecture

main.py            — FastAPI app, model loading, streaming TTS endpoint
download_model.py  — Downloads model + voice pack on first start
entrypoint.sh      — Runs download then starts uvicorn
Dockerfile         — CUDA 12.8 runtime image (Ada Lovelace + Blackwell)
benchmark/         — Benchmark suite with Whisper validation

⁠Requirements

  • NVIDIA GPU with CUDA 12.8+ (RTX 4090 / RTX 5090 and similar)
  • Docker with NVIDIA Container Toolkit

Tag summary

Content type

Image

Digest

sha256:e893b4f5c…

Size

6.9 GB

Last updated

5 months ago

docker pull ffaerber/kokoro-cuda