OpenAI-compatible text-to-speech HTTP API wrapping kyutai-labs/pocket-tts.
146
OpenAI-compatible text-to-speech HTTP API wrapping kyutai-labs/pocket-tts. The model is loaded once at container startup and kept warm in memory for every request. Model weights are baked into the image at build time (offline at runtime) for a fast, predictable startup.
pre-built with english (alba) and french (estelle, azelma)
Project version uses semver Major.Minor.Patch.
GET /health{"status": "ok", "model_loaded": true, "sample_rate": 24000}
POST /v1/audio/speechSame request/response contract as kokoro-fastapi's /v1/audio/speech, so this image is a drop-in replacement for any client already speaking that contract.
curl -X POST http://localhost:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"model": "pocket-tts", "voice": "alba", "input": "Hello, world.", "response_format": "wav"}' \
-o out.wav
response_format accepts wav, pcm (no extra dependency), or mp3/opus/aac/flac (requires the image to be built with ffmpeg, on by default).
| Variable | Default | Effect |
|---|---|---|
HOST / PORT | 0.0.0.0 / 8000 | uvicorn bind address |
POCKET_TTS_LANGUAGE | english | model language |
POCKET_TTS_CONFIG | (empty) | custom YAML config path/URL/hf:// |
POCKET_TTS_CHECKPOINT | (empty) | custom training checkpoint |
POCKET_TTS_DEVICE | cpu | cpu or a CUDA device string |
POCKET_TTS_QUANTIZE | false | requires the image built with --build-arg ENABLE_QUANTIZE=true |
POCKET_TTS_TEMPERATURE | 0.7 | sampling temperature |
POCKET_TTS_SAMPLER_DECODE_STEPS | 1 | sampler decode steps |
POCKET_TTS_NOISE_CLAMP | (empty) | noise clamp value |
POCKET_TTS_EOS_THRESHOLD | -4.0 | end-of-speech threshold |
POCKET_TTS_LSD_DECODE_STEPS | (empty) | advanced sampler param, passed through only if set |
POCKET_TTS_MAX_TOKENS | (empty, library default) | max tokens per generation |
POCKET_TTS_FRAMES_AFTER_EOS | (empty, library default) | frames generated after EOS |
POCKET_TTS_VOICE | alba | default voice when a request omits one |
POCKET_TTS_ALLOW_ARBITRARY_VOICE | false | security-sensitive. If false (default), only catalog voice names are accepted in requests. If true, a request's voice can be an arbitrary local path or hf:// URL, which pocket-tts will read/download — only enable this if callers of the API are trusted, since it allows local file access and outbound requests to be triggered from the request body. |
POCKET_TTS_DEFAULT_RESPONSE_FORMAT | wav | default response format |
POCKET_TTS_NUM_THREADS | 2 | torch.set_num_threads() |
POCKET_TTS_MAX_CONCURRENCY | 1 | max concurrent generations (model is not thread-safe) |
POCKET_TTS_MAX_INPUT_CHARS | 5000 | rejects oversized input |
POCKET_TTS_API_KEY | (empty, disabled) | if set, requires Authorization: Bearer <key> |
Content type
Image
Digest
sha256:307cc7e44…
Size
2 GB
Last updated
about 1 month ago
docker pull epidoux/pocket-tts:0.1.1