Sign inSign up

epidoux/pocket-tts

By epidoux

•Updated about 1 month ago

OpenAI-compatible text-to-speech HTTP API wrapping kyutai-labs/pocket-tts.

Image
Machine learning & AI
0

146

epidoux/pocket-tts repository overview

⁠Docker image for pocket-tts

OpenAI-compatible text-to-speech HTTP API wrapping kyutai-labs/pocket-tts⁠. The model is loaded once at container startup and kept warm in memory for every request. Model weights are baked into the image at build time (offline at runtime) for a fast, predictable startup.

pre-built with english (alba) and french (estelle, azelma)

Project version uses semver Major.Minor.Patch.

⁠API

⁠GET /health
{"status": "ok", "model_loaded": true, "sample_rate": 24000}
⁠POST /v1/audio/speech

Same request/response contract as kokoro-fastapi⁠'s /v1/audio/speech, so this image is a drop-in replacement for any client already speaking that contract.

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model": "pocket-tts", "voice": "alba", "input": "Hello, world.", "response_format": "wav"}' \
  -o out.wav

response_format accepts wav, pcm (no extra dependency), or mp3/opus/aac/flac (requires the image to be built with ffmpeg, on by default).

⁠Environment variables

VariableDefaultEffect
HOST / PORT0.0.0.0 / 8000uvicorn bind address
POCKET_TTS_LANGUAGEenglishmodel language
POCKET_TTS_CONFIG(empty)custom YAML config path/URL/hf://
POCKET_TTS_CHECKPOINT(empty)custom training checkpoint
POCKET_TTS_DEVICEcpucpu or a CUDA device string
POCKET_TTS_QUANTIZEfalserequires the image built with --build-arg ENABLE_QUANTIZE=true
POCKET_TTS_TEMPERATURE0.7sampling temperature
POCKET_TTS_SAMPLER_DECODE_STEPS1sampler decode steps
POCKET_TTS_NOISE_CLAMP(empty)noise clamp value
POCKET_TTS_EOS_THRESHOLD-4.0end-of-speech threshold
POCKET_TTS_LSD_DECODE_STEPS(empty)advanced sampler param, passed through only if set
POCKET_TTS_MAX_TOKENS(empty, library default)max tokens per generation
POCKET_TTS_FRAMES_AFTER_EOS(empty, library default)frames generated after EOS
POCKET_TTS_VOICEalbadefault voice when a request omits one
POCKET_TTS_ALLOW_ARBITRARY_VOICEfalsesecurity-sensitive. If false (default), only catalog voice names are accepted in requests. If true, a request's voice can be an arbitrary local path or hf:// URL, which pocket-tts will read/download — only enable this if callers of the API are trusted, since it allows local file access and outbound requests to be triggered from the request body.
POCKET_TTS_DEFAULT_RESPONSE_FORMATwavdefault response format
POCKET_TTS_NUM_THREADS2torch.set_num_threads()
POCKET_TTS_MAX_CONCURRENCY1max concurrent generations (model is not thread-safe)
POCKET_TTS_MAX_INPUT_CHARS5000rejects oversized input
POCKET_TTS_API_KEY(empty, disabled)if set, requires Authorization: Bearer <key>

Tag summary

Content type

Image

Digest

sha256:307cc7e44…

Size

2 GB

Last updated

about 1 month ago

docker pull epidoux/pocket-tts:0.1.1