Sign inSign up

stormotron/stormradio-vdj

By stormotron

•Updated 26 days ago

A lightweight, Docker-ready virtual DJ engine for a radio station, written in Python.

Image
0

56

stormotron/stormradio-vdj repository overview

⁠StormRadio-vDJ, v0.0.3

A lightweight, Docker-ready virtual DJ engine for a radio station, written in Python. StormRadio-vDJ turns "now playing" metadata into a short, natural-sounding spoken announcement: it asks an LLM to write a DJ line naming the current and next track, synthesizes it with a TTS engine, and returns a ready-to-stream MP3 (128 kbit/s) over a simple authenticated HTTP API. It has no database, no external volumes, and no persistent state - every request is self-contained.

⁠Directory Structure

.
├── Dockerfile
├── docker-compose.yml
├── requirements.txt
├── main.py              # entry point: startup healthchecks, then starts the server
├── app/
│   ├── config.py         # reads environment variables, holds defaults
│   ├── version.py         # product name & version, shared by /health and startup logs
│   ├── prompts.py         # hardcoded per-language DJ-announcement prompts, selected via AI_LANG
│   ├── logging_setup.py   # unified, thread-aware logger
│   ├── executors.py       # bounded thread pools for the AI and TTS engines
│   ├── audio.py           # PCM<->MP3 conversion via av/PyAV (no external tools)
│   ├── errors.py          # short, traceback-free messages for auth (401/403) failures
│   ├── server.py          # aiohttp application, authorization, routes
│   └── engines/
│       ├── base.py         # AIEngine / TTSEngine abstract interfaces
│       ├── ai_groq.py       # AI engine on top of Groq
│       ├── ai_gemini.py     # AI engine on top of the Gemini API (text generation)
│       ├── tts_edge.py      # TTS engine on top of edge-tts, returns mp3 128 kbit/s
│       └── tts_gemini.py    # TTS engine on top of the Gemini API (speech generation), returns mp3 128 kbit/s
└── README.md

⁠Quick Start

⁠1. Prerequisites
  • Docker⁠ installed on your host system.
  • Depending on the chosen AI engine: a Groq⁠ API key for groq (default), or a Gemini API key⁠ for gemini.
  • Depending on the chosen TTS engine: nothing extra for edge_tts (default), or a Gemini API key⁠ for gemini_tts.
⁠2. Build & Run
# edit docker-compose.yml directly: set GROQ_API_KEY and TOKEN under
# environment:, plus GEMINI_TTS_API_KEY if you'll use gemini_tts
docker compose up --build

Or directly with Docker:

docker build -t stormradio-vdj .

docker run --rm -p 8080:8080 \
  -e GROQ_API_KEY=your_groq_api_key \
  -e TOKEN=your_32_byte_token \
  stormradio-vdj

On startup the app initializes the configured AI and TTS engines and immediately runs a healthcheck for each of them (a test prompt for the AI engine, a test phrase for the TTS engine), measuring latency. If either check fails, the process does not exit - it halts and waits forever without accepting any clients, so it won't get caught in a restart loop - see Startup healthchecks below.

⁠3. Usage
⁠Requesting an announcement
curl -X POST http://localhost:8080/get_back_forward \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "AI_TRACK_INFO": "Imagine Dragons - Believer",
        "AI_NEXT_TRACK_INFO": "Coldplay - Yellow",
        "AI_LANG": "en",
        "AI_DJ_NAME": "Max"
      }' \
  --output announcement.mp3

See API below for the full request/response contract.

⁠Health/version check
curl http://localhost:8080/health \
  -H "Authorization: Bearer $TOKEN"
{"product": "StormRadio-vDJ", "version": "0.0.3", "status": "ok"}

⁠Environment Variables

VariableDefaultDescription
AI_ENGINEgroqText-generation engine. groq or gemini.
AI_ENGINE_MAX_THREADS1Max concurrent requests to the AI engine; further requests queue instead of running in parallel.
TTS_ENGINEedge_ttsSpeech-synthesis engine. edge_tts or gemini_tts.
TTS_ENGINE_MAX_THREADS1Max concurrent requests to the TTS engine; further requests queue instead of running in parallel.
GROQ_API_KEY-Groq API key. Required if AI_ENGINE=groq.
TOKEN-Client authorization token, sent by clients as Authorization: Bearer <TOKEN>. Required on every request to /get_back_forward and /health. Should be at least 32 bytes. Required.
GROQ_MODELopenai/gpt-oss-120bGroq model used for generation. Groq periodically retires models - if the AI healthcheck fails with model_not_found, check https://console.groq.com/docs/models⁠ and override this.
GEMINI_API_KEY-Gemini API key from https://aistudio.google.com/apikey⁠, used by the AI engine. Required if AI_ENGINE=gemini. Independent from GEMINI_TTS_API_KEY - set both if you use Gemini for both text and speech, even to the same key value.
GEMINI_MODELgemini-2.5-flashGemini model used for text generation. Used only when AI_ENGINE=gemini.
EDGE_TTS_VOICEru-RU-DmitryNeuraledge-tts voice. Used only when TTS_ENGINE=edge_tts.
GEMINI_TTS_API_KEY-Gemini API key from https://aistudio.google.com/apikey⁠. Required if TTS_ENGINE=gemini_tts.
GEMINI_TTS_MODELgemini-2.5-flash-preview-ttsGemini TTS model. See https://ai.google.dev/gemini-api/docs/speech-generation⁠ for current model names. Used only when TTS_ENGINE=gemini_tts.
GEMINI_TTS_VOICEKoreOne of the prebuilt Gemini TTS voice names. Used only when TTS_ENGINE=gemini_tts.
HOST0.0.0.0Bind address for the HTTP server.
PORT8080Bind port for the HTTP server.
SOUND_START_SILENCE_PADDING2Seconds of silence prepended to the returned MP3, before the spoken announcement.
SOUND_END_SILENCE_PADDING2Seconds of silence appended to the returned MP3, after the spoken announcement.

All defaults live in app/config.py, so the Dockerfile intentionally does not hardcode any of them - they are meant to be supplied by whoever runs the container (via docker run -e ... or docker-compose.yml), and the code already falls back to sensible defaults when a variable is not set.

AI_ENGINE and TTS_ENGINE are designed as extension points - support for local models and other TTS providers is planned for future versions.

⁠API

⁠POST /get_back_forward

Generates a DJ announcement for the current/next track and returns it as an MP3 stream.

Headers:

Authorization: Bearer <TOKEN>
Content-Type: application/json

Body:

FieldRequiredDescription
AI_TRACK_INFOyesThe currently playing track, e.g. "Artist - Title".
AI_NEXT_TRACK_INFOyesThe next track to be announced, e.g. "Artist - Title".
AI_LANGnoru, en, zh, es, fr, or de. Selects which of the built-in announcement prompts (see app/prompts.py) is sent to the AI engine. Defaults to en. Any other value returns 400.
AI_DJ_NAMEnoThe DJ's on-air name, inserted into the prompt. Defaults to vDJ. Whitespace/newlines are collapsed and the value is capped at 64 characters.

The prompts are fixed in app/prompts.py, one per supported AI_LANG (ru, en, zh, es, fr, de), built with AI_DJ_NAME substituted in. Each prompt asks for a natural-sounding, 5–8 second DJ line that names the current and next track, and instructs the model to sometimes - not on every call - open with a brief self-introduction (e.g. "This is DJ vDJ, and right now we're playing...") instead of always announcing directly, so back-to-back announcements don't sound identically scripted.

Responses:

StatusMeaning
200Body is a binary MP3 stream (audio/mpeg, 128 kbit/s).
400AI_TRACK_INFO/AI_NEXT_TRACK_INFO missing, or AI_LANG is not one of ru, en, zh, es, fr, de.
401Missing or incorrect Authorization token.
500AI/TTS engine failure, audio conversion failure, or any other unexpected error.

Example:

curl -X POST http://localhost:8080/get_back_forward \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "AI_TRACK_INFO": "Imagine Dragons - Believer",
        "AI_NEXT_TRACK_INFO": "Coldplay - Yellow",
        "AI_LANG": "en",
        "AI_DJ_NAME": "Max"
      }' \
  --output announcement.mp3
⁠GET /health

Simple liveness/version endpoint. Requires the same Authorization: Bearer <TOKEN> header as /get_back_forward; a missing or incorrect token returns 401.

{"product": "StormRadio-vDJ", "version": "0.0.3", "status": "ok"}

version is read from app/version.py (APP_VERSION), the same constant logged at startup - it isn't hand-duplicated between the code and this document.

⁠Startup healthchecks

Before accepting any client, the app runs one request through each configured engine and measures latency:

  • AI engine: sends "Reply to the test message" and expects a text completion.
  • TTS engine: synthesizes "Say hello to me" and expects non-empty audio.

If either fails (including a missing TOKEN, or the AI/TTS engine failing to even initialize), the process does not exit. It logs the failure and then halts, blocking forever without ever starting the HTTP server or accepting a client. This is intentional: a container under restart: unless-stopped (or any orchestrator that retries on exit) would otherwise crash-loop on a persistent problem like a bad API key, spamming the same error over and over. Instead, the very last line of the log is always:

stormradio-vdj has HALTED due to a startup error above and will NOT restart automatically - fix the problem and restart this container manually (e.g. `docker restart <container>`)

Fix whatever the error above it says (missing/invalid TOKEN, GROQ_API_KEY, GEMINI_API_KEY, GEMINI_TTS_API_KEY, an unreachable model, etc.) and restart the container yourself - docker restart <container> or docker compose up -d --force-recreate.

Authorization failures (401/403 from Groq or Gemini - wrong or missing API key) are logged as a single short line via app/errors.py instead of a full stack trace; any other failure (network issues, model not found, unexpected errors) is logged with the full traceback for debugging.

⁠Concurrency model

Every request to /get_back_forward is handled asynchronously by the aiohttp event loop, so the server itself never blocks on I/O. Inside the handler:

  • the fixed prompt (per AI_LANG) and track metadata are sent to the AI engine through a bounded thread pool (AI_ENGINE_MAX_THREADS) - if all AI workers are busy, the request simply waits in that pool's queue;
  • the generated announcement text is sent to the TTS engine through its own bounded thread pool (TTS_ENGINE_MAX_THREADS);
  • the resulting audio goes through the shared MP3 conversion step described below;
  • the result is returned to the client as audio/mpeg.

Logging is unified across the whole application - every log line shows which thread it came from, so parallel request processing stays observable, and each request is tagged with a short request ID for correlating its log lines end to end.

⁠MP3 conversion without external tools

The application never spawns ffmpeg or any other external program via a subprocess. Decoding and encoding MP3 happens entirely in-process through the av package (PyAV) - Python bindings to the FFmpeg libraries (libavcodec, libavformat, libswresample), statically linked into the package itself. PyPI publishes prebuilt wheels for av, including musllinux-compatible ones (for Alpine), so pip install requires no compilation, no CMake, no compilers, and no network access beyond PyPI itself. At runtime, no external binaries are ever executed - only function calls inside app/audio.py.

Both TTS engines feed into the same conversion step (app/audio.py), regardless of what format they natively return:

  • edge_tts returns an already-compressed MP3 (24 kHz / mono / 48 kbit/s) - it's decoded back to PCM and re-encoded once, to a guaranteed 128 kbit/s.
  • gemini_tts returns raw 16-bit PCM at 24 kHz mono directly, with no intermediate lossy step - it's encoded straight to 128 kbit/s.

Both engines funnel into the same pcm_to_mp3() in app/audio.py, and that's the single place silence padding is applied: before encoding, SOUND_START_SILENCE_PADDING seconds of silence are prepended and SOUND_END_SILENCE_PADDING seconds are appended to the raw PCM (2 seconds each by default), lengthening the resulting MP3 accordingly. Since it lives in the shared conversion step rather than in each engine, it applies uniformly regardless of TTS_ENGINE and any future engine gets it for free.

⁠Choosing a TTS engine

edge_tts (default) is an unofficial wrapper around the "read aloud" feature of the Edge browser. It requires no API key and works out of the box, but Microsoft always returns audio at 24 kHz / mono / 48 kbit/s, and since that source is already lossy before the app's own re-encode, the perceived quality ceiling is set by that low bitrate, not by the final 128 kbit/s output.

gemini_tts uses Google's Gemini API speech generation for noticeably more natural prosody:

  1. Get a free API key from https://aistudio.google.com/apikey⁠ - no GCP project, no service account, no IAM role to configure.
  2. Set TTS_ENGINE=gemini_tts and GEMINI_TTS_API_KEY.
  3. Optionally override GEMINI_TTS_MODEL and GEMINI_TTS_VOICE (see the Environment Variables table).

Note: Gemini TTS models are labeled "Preview" by Google - the model lineup moves fast (e.g. gemini-2.5-flash-preview-tts has already been followed by gemini-3.1-flash-tts-preview), rate limits are tighter than stable models, and pricing is billed per output audio token rather than having a guaranteed permanent free tier like some other Gemini models. Check https://ai.google.dev/gemini-api/docs/pricing⁠ for current terms before relying on this in production, and adjust GEMINI_TTS_MODEL if Google retires the default.

Known issue: Google's preview TTS models occasionally return HTTP 200 with an empty candidate and finish_reason=OTHER - no audio, no error, unrelated to which voice you picked (this is a documented, intermittent bug on Google's side, not a sign of an invalid GEMINI_TTS_VOICE). tts_gemini.py retries such empty responses up to 3 times before giving up with a clear error message; if it still fails consistently, try again shortly, try a different GEMINI_TTS_VOICE, or switch to TTS_ENGINE=edge_tts.

Available prebuilt voices (GEMINI_TTS_VOICE) include, among others: Zephyr, Puck, Charon, Kore (default), Fenrir, Leda, Orus, Aoede - see https://ai.google.dev/gemini-api/docs/speech-generation⁠ for the full list of 30 voices and their character descriptions (e.g. Puck is "Upbeat", Charon is "Informative").

⁠Extending the engines

To add a new AI or TTS engine:

  1. Implement a class inheriting from AIEngine or TTSEngine in app/engines/base.py.
  2. Register it in build_ai_engine / build_tts_engine in main.py, keyed off the corresponding environment variable value.

The bounded thread pools (AI_ENGINE_MAX_THREADS / TTS_ENGINE_MAX_THREADS) work the same way for any engine - the limit is enforced at the ThreadPoolExecutor level, not inside the engine itself, and the async HTTP handler simply awaits loop.run_in_executor(...) on the appropriate pool.

⁠Security notes

  • Every call to /get_back_forward and /health requires Authorization: Bearer <TOKEN>; a missing or incorrect token returns 401 without touching the AI/TTS engines. /health is intentionally not exempt, since it also returns product/version information.
  • TOKEN is compared as a plain string - generate it with something like openssl rand -hex 32 and treat it as a secret, the same as GROQ_API_KEY / GEMINI_API_KEY / GEMINI_TTS_API_KEY.
  • The app has no database, no session state, and no mounted volumes - nothing about a request is persisted beyond the lifetime of handling it.
  • Authorization failures against Groq/Gemini (bad API key) are logged as a short message, not a full traceback - see Startup healthchecks above.

Tag summary

Content type

Image

Digest

sha256:2566dd683…

Size

77 MB

Last updated

26 days ago

docker pull stormotron/stormradio-vdj