Thirty audio engines. One port. Zero cloud. Fire-and-forget async jobs. Webhooks.
You needed Demucs for stems. Then librosa for BPM and key. Then basic-pitch for MIDI transcription. Then pyannote for speaker diarization. Then DeepFilterNet for speech enhancement. Then you spent three days debugging Python version conflicts and now you hate everything.
audiolla is what happens when you stop doing that.
Every audio processing tool worth using β wrapped in one HTTP API, running in one Docker container. POST a file. Get audio, JSON, or MIDI back. Drive it from curl, shell scripts, Python notebooks, Makefiles, or point an LLM agent at the MCP endpoint and let it rip.
No account. No subscription. No per-minute billing. No vendor lock-in. docker run and you're done.
| ποΈ Stem separation | Demucs β htdemucs, fine-tuned, 6-stem, MDX variants |
| ποΈ Mastering | Reference mastering (matchering) + custom pedalboard chains |
| π Analysis | BPM Β· key Β· LUFS Β· beats Β· onsets Β· melody Β· structural segments |
| πΉ Chords + key | Chord detection + Krumhansl-Schmuckler key estimation |
| π΅ Audio β MIDI | Polyphonic transcription via Spotify's basic-pitch (ONNX, no TF) |
| π§Ή Restoration | De-reverb Β· de-echo Β· de-noise via UVR BS-Roformer + MelBand Roformer |
| π£οΈ Speech | Enhancement (DeepFilterNet) Β· VAD (silero-vad) Β· diarization (pyannote) |
| πΌοΈ Visuals | Spectrogram + waveform PNGs + 8-mode animated MP4/WebM |
| π Fingerprint | Chromaprint acoustic fingerprinting (AcoustID-compatible) |
| βοΈ Silence | Detect gaps Β· trim edges Β· strip all silence |
| πΌ MIDI pipeline | Compose from JSON Β· inspect Β· transform Β· render via fluidsynth |
| πΈ Effects | 23-effect pedalboard chain β Compressor, Reverb, PitchShift, filtersβ¦ |
| π§ Transforms | Sox DSP β pitch, tempo, EQ, reverb, gain |
| π’ Loudness | Measure LUFS Β· normalize to target |
| π₯ HPSS | Harmonic/percussive source separation via librosa median filter |
| π Noise reduction | Spectral noise reduction via noisereduce β stationary + adaptive modes |
| β© Time-stretch | Independent tempo factor + pitch shift via librosa phase vocoder |
| π·οΈ Audio tagging | Top-K AudioSet class labels via Audio Spectrogram Transformer |
| π Audio embeddings | 512-dim semantic embeddings via LAION CLAP + optional text similarity |
| π·οΈ Zero-shot classify | CLAP cosine similarity against any free-form text labels β genres, moods, instruments |
| π Audio info | ffprobe metadata β duration, sample rate, channels, codec, bit depth |
| βοΈ Trim | Cut a clip by start/end seconds β any format in, any format out |
| ποΈ Mix | Combine N staged tracks with per-track gain_db β pure ffmpeg, no model |
| π Concat | Stitch N audio files end-to-end in order |
| β© Speed | Change playback speed without pitch shift (0.1Γ β 10Γ) via ffmpeg atempo |
| π Convert | Re-encode: format, sample rate, channel count in one call |
| π Similar | Cosine similarity between two audio files via CLAP embeddings |
| πΉ MIDI quantize | Snap MIDI note timings to a rhythmic grid (16th, 8th, quarterβ¦) |
| π Fade | Fade-in and/or fade-out with 13 curve shapes |
| βͺ Reverse | Flip audio backwards |
| π Loop | Repeat audio N times |
| π― BPM match | Auto-detect BPM then stretch to a target β no manual math |
| π Loudness curve | RMS envelope over time β time-stamped dB values for gain automation |
| π€ Pitch correct | Auto-tune toward nearest chromatic semitone β configurable strength |
| π§ Repair | Declip + dehum β fix clipped peaks and remove power-line hum |
| π Loop point | Find best seamless loop boundary β score, bar count, candidates list |
| π₯ Drum machine | Step-sequencer spec β GM drum MIDI β 16-step pattern, swing, tempo |
| πΌ Chords to MIDI | Chord progression β MIDI file β root+3rd+5th voicings per segment |
| βοΈ Stereo width | Widen or collapse the stereo image via M/S processing |
| βοΈ Split | Split into N equal parts or on silence β returns ZIP of segments |
| π Pan | Position audio in the stereo field (-1 left β 0 center β 1 right) |
| ποΈ EQ | Parametric EQ β JSON array of freq/gain_db/width_hz bands |
| π΅ Key match | Detect source key then pitch-shift to a target key |
| ποΈ Sidechain duck | Duck music when a trigger track (voice) is loud |
| π·οΈ Metadata | Read and write ID3/Vorbis/FLAC/WAV audio tags via mutagen |
| π΄ Clip detect | Detect digital clipping β count, ratio, peak dBFS |
| βοΈ Mid/Side | Encode L/R β Mid+Side or decode Mid+Side β L/R |
| βοΈ Beat slice | Slice audio at detected beat positions β returns ZIP of segments |
| ποΈ Conv reverb | Convolution reverb via impulse response β wet_mix control |
| π₯ Transient shaper | Attack/sustain dual-compressor β punch up drums, cut room tail |
| ποΈ Multiband compress | N-band compressor with zero-phase LR4 crossovers β mastering-grade dynamics |
| ποΈ DJ prep | One call: BPM + key + Camelot wheel position + integrated LUFS |
| π¦ Batch | Run trim/convert/fade/reverse/speed/eq on staged files in sequence |
| π§© Presets + pipeline | Curated YAML workflows (master-for-spotify, podcast-cleanup, β¦) + ad-hoc op chaining server-side |
| ποΈ Catalog | GET /v1/catalog β machine-readable endpoint list grouped by category for discovery |
| β‘ Async jobs | Every endpoint supports async_job=true β fire-and-forget + webhook callbacks |
# no GPU
docker run --rm -it \
-v $HOME/.audiolla-data:/data \
-p 8000:8000 \
psyb0t/audiolla:latest
# GPU
docker run --rm -it --gpus all \
-v $HOME/.audiolla-data:/data \
-e AUDIOLLA_DEVICE=cuda \
-p 8000:8000 \
psyb0t/audiolla:latest-cuda
Demucs weights prefetch at container startup (for whichever variants are enabled) and cache in /data/torch_cache/. First boot downloads them; same -v mount next time and they're already there. Other engines (matchering, pedalboard, librosa, sox, fx, midi) have no weights β they're ready as soon as /healthz is green.
v1.0.0 is a breaking API release. Every existing client breaks. The new shape:
multipart/form-data except at /v1/files)file_path (FILES_DIR-relative) xor file_url (server-side fetch). Pre-stage the file via PUT /v1/files/{path} first.output_path xor output_url. No more raw audio bytes in responses.async_job=true auto-stages to jobs/{id}.{ext} if neither output is given.audio_base64 (and midi_base64 / image_base64 / video_base64). Same output_path xor output_url requirement.openapi.yaml is now the contract β Pydantic models regenerate from it via make generate. Never hand-edit src/audiolla/schema/_generated.py.- curl -X POST http://localhost:8000/v1/audio/normalize \
- -F "[email protected]" -F "target_lufs=-14" -o normalized.wav
+ # 1) stage the file (multipart only lives here now)
+ curl -X PUT --data-binary @track.wav \
+ -H 'Content-Type: application/octet-stream' \
+ http://localhost:8000/v1/files/uploads/track.wav
+ # 2) process via JSON body β response is JSON, not bytes
+ curl -X POST http://localhost:8000/v1/audio/normalize \
+ -H 'Content-Type: application/json' \
+ -d '{"file_path":"uploads/track.wav","target_lufs":-14,"output_path":"out/normalized.wav"}'
+ # 3) retrieve the result
+ curl -o normalized.wav http://localhost:8000/v1/files/out/normalized.wav
Why? See the v1.0.0 CHANGELOG entryβ for the full rationale.
Once the container is up, this is a complete audio pipeline in six commands (every audio endpoint is JSON-body now; stage your input file at /v1/files/... first):
# stage your input file
curl -X PUT --data-binary @song.wav \
-H 'Content-Type: application/octet-stream' \
http://localhost:8000/v1/files/uploads/song.wav
# rip the vocals out of a track
curl -X POST http://localhost:8000/v1/audio/separate \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/song.wav","engine":"htdemucs","stems":["vocals"],"output_path":"out/vocals.wav"}'
# β {"path":"out/vocals.wav","size":...,"output_format":"wav"}
curl -o vocals.wav http://localhost:8000/v1/files/out/vocals.wav
# what key is it in? what are the chords?
curl -X POST http://localhost:8000/v1/audio/chords \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/song.wav"}'
# β {"key":"F# minor","key_confidence":0.91,"chords":[{"chord":"F#m","start_sec":0.0,...},...]}
# transcribe that vocal melody to MIDI
curl -X PUT --data-binary @out/vocals.wav -H 'Content-Type: application/octet-stream' \
http://localhost:8000/v1/files/uploads/vocals.wav # only if not already staged
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/vocals.wav","output_path":"out/melody.mid"}'
# render the MIDI back to audio through a SoundFont
curl -X POST http://localhost:8000/v1/midi/render \
-H 'Content-Type: application/json' \
-d '{"file_path":"out/melody.mid","output_path":"out/rendered.wav"}'
# strip background noise from a voice recording
curl -X POST http://localhost:8000/v1/audio/noise-reduce/uvr-denoise \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/interview.wav","output_path":"out/clean.wav"}'
# who's speaking and when?
curl -X POST http://localhost:8000/v1/audio/diarize/pyannote \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/interview.wav"}'
# β {"num_speakers":2,"segments":[{"speaker":"SPEAKER_00","start_sec":0.5,"end_sec":8.2},...]}
Audio in. MIDI out. Chords detected. Speakers identified. De-noised. Re-synthesized. No Python environment to set up. No API keys. No account. Just HTTP.
Output defaults to wav. Add "output_format":"mp3" to the JSON body to get mp3 instead (flac, opus, aac, pcm also work).
Every audio endpoint takes an application/json body. The only place multipart still lives is PUT /v1/files/{path} (raw bytes for staging an input file).
Input β every audio endpoint requires exactly one of:
file_path β path inside the /v1/files staging area (stage with PUT /v1/files/{path} first)file_url β remote URL the server fetches (disabled by default β see Remote URLsβ )Output β audio-producing endpoints require exactly one of:
output_path β server writes to /v1/files/<path>, returns JSON {"path":..., "size":..., ...}output_url β server PUTs to a presigned URL, returns JSON {"url":..., "size":..., ...}Analysis-only endpoints (those that return JSON data, e.g. /v1/audio/analyze, /v1/audio/loudness, /v1/audio/info) don't need output_path / output_url β the response is the result.
# stage input
curl -X PUT --data-binary @track.wav \
-H 'Content-Type: application/octet-stream' \
http://localhost:8000/v1/files/uploads/track.wav
# vocals only
curl -X POST http://localhost:8000/v1/audio/separate \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","engine":"htdemucs","stems":["vocals"],"output_path":"out/vocals.wav"}'
curl -o vocals.wav http://localhost:8000/v1/files/out/vocals.wav
# all 4 stems as a ZIP
curl -X POST http://localhost:8000/v1/audio/separate \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","engine":"htdemucs","output_path":"out/stems.zip"}'
curl -o stems.zip http://localhost:8000/v1/files/out/stems.zip
# stage track + reference
curl -X PUT --data-binary @track.wav -H 'Content-Type: application/octet-stream' \
http://localhost:8000/v1/files/uploads/track.wav
curl -X PUT --data-binary @ref.wav -H 'Content-Type: application/octet-stream' \
http://localhost:8000/v1/files/uploads/ref.wav
# match EQ + loudness to a reference track
curl -X POST http://localhost:8000/v1/audio/master \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","mode":"reference","reference_path":"uploads/ref.wav","output_path":"out/mastered.wav"}'
curl -o mastered.wav http://localhost:8000/v1/files/out/mastered.wav
# run a built-in pedalboard chain (presets: transparent, loud)
curl -X POST http://localhost:8000/v1/audio/master \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","mode":"chain","preset":"loud","output_path":"out/mastered.wav"}'
curl -o mastered.wav http://localhost:8000/v1/files/out/mastered.wav
# returns JSON. features: bpm, key, loudness, duration,
# spectral_centroid, rms, zcr. Omit features to get them all.
curl -X POST http://localhost:8000/v1/audio/analyze \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","features":["bpm","key","loudness"]}'
# beat grid β returns bpm + beat timestamps
curl -X POST http://localhost:8000/v1/audio/beats \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav"}'
# onset timestamps β note attacks, transients
curl -X POST http://localhost:8000/v1/audio/onsets \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav"}'
# dominant melody contour β pitch in Hz per frame
curl -X POST http://localhost:8000/v1/audio/melody \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav"}'
# structural segmentation β labels recurring sections A, B, C...
curl -X POST http://localhost:8000/v1/audio/segments \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","num_segments":6}'
Beat detection also generates a click-track file when click_track=true (set output_path to receive it) β handy for aligning a mix to a grid. Pass start_bpm=140 to seed the tracker when you already know the rough tempo (faster, more accurate). Melody can be exported as a single-track MIDI file via as_midi=true + output_path.
# find silent gaps in a recording (no trim_mode β JSON only)
curl -X POST http://localhost:8000/v1/audio/silence \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","threshold_db":-30,"min_duration_sec":1.0}'
# trim all silence and stage the result
curl -X POST http://localhost:8000/v1/audio/silence \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","threshold_db":-30,"min_duration_sec":0.5,"trim_mode":"all","output_path":"out/trimmed.wav"}'
curl -o trimmed.wav http://localhost:8000/v1/files/out/trimmed.wav
# trim only leading/trailing silence
curl -X POST http://localhost:8000/v1/audio/silence \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","threshold_db":-40,"min_duration_sec":0.3,"trim_mode":"edges","output_path":"processed/trimmed.wav"}'
trim_mode=edges β chop leading + trailing silence only. trim_mode=all β remove every detected gap (compress a talk recording, tighten a loop). Without trim_mode, the response is JSON only: silent_ranges, non_silent_ranges, duration β and output_path / output_url is not required.
Visual output splits into two sub-namespaces by output type:
# Static PNG spectrogram (color + scale params)
curl -X POST http://localhost:8000/v1/audio/visualize/image/spectrogram \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","width":1280,"height":720,"output_path":"out/spec.png"}'
curl -o spec.png http://localhost:8000/v1/files/out/spec.png
# Static PNG waveform (color param)
curl -X POST http://localhost:8000/v1/audio/visualize/image/waveform \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","width":1280,"height":240,"output_path":"out/wave.png"}'
curl -o wave.png http://localhost:8000/v1/files/out/wave.png
# Animated MP4 spectrum analyser (fps + container params)
curl -X POST http://localhost:8000/v1/audio/visualize/video/spectrum \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","width":1280,"height":720,"fps":30,"container":"mp4","output_path":"out/viz.mp4"}'
curl -o viz.mp4 http://localhost:8000/v1/files/out/viz.mp4
/image/spectrogram: produces a PNG (staged via output_path or PUT to output_url). Params: width, height, color (default intensity), scale (log/lin).
/image/waveform: produces a PNG. Params: width, height, color (default lime).
/video/{mode}: spectrum (scrolling FFT), waves (oscilloscope), cqt (constant-Q transform), freqs (bar-graph analyzer), volume (VU meter), vectorscope (stereo X/Y scope), phasemeter, histogram. Params: width, height, fps, container (mp4 default, webm).
# Chromaprint fingerprint β identifies a recording regardless of encoding
curl -X POST http://localhost:8000/v1/audio/fingerprint \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav"}'
# β {"duration": 215.34, "fingerprint": "AQADtEqRRIuQ..."}
# include the raw integer array (for custom similarity scoring)
curl -X POST http://localhost:8000/v1/audio/fingerprint \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","return_raw":true}'
The base64 fingerprint string is compatible with the AcoustIDβ lookup service.
AI audio restoration via UVR ecosystem models β BS-Roformer and MelBand Roformer. All three are unified under POST /v1/audio/restore/{engine}.
# Remove room reverb (BS-Roformer, SDR 19+)
curl -X POST http://localhost:8000/v1/audio/restore/uvr-dereverb \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","output_path":"out/dry.wav"}'
# Remove echo β normal mode
curl -X POST http://localhost:8000/v1/audio/restore/uvr-deecho \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","output_path":"out/noecho.wav"}'
# Remove echo β aggressive mode (same engine, harder suppression)
curl -X POST http://localhost:8000/v1/audio/restore/uvr-deecho \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","aggressive":true,"output_path":"out/noecho.wav"}'
# Remove broadband background noise β ML (MelBand Roformer, SDR 28)
curl -X POST http://localhost:8000/v1/audio/restore/uvr-denoise \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/track.wav","output_path":"out/clean.wav"}'
All support output_format, output_path, output_url. For DSP-based noise reduction (no GPU) use noise-reduce/noise-reduce.
UVR engines also work through /v1/audio/separate β uvr-vocal-bsr (BS-Roformer, SDR 13) and uvr-karaoke return vocal + instrumental stems like Demucs but often with higher quality.
Polyphonic audio-to-MIDI via Spotify's basic-pitch (ONNX backend, no TensorFlow). Play guitar, hum a melody, record a piano riff β get a MIDI file back with all the notes.
# Any audio β MIDI file (staged)
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/guitar_riff.wav","output_path":"out/riff.mid"}'
curl -o riff.mid http://localhost:8000/v1/files/out/riff.mid
# Tune the detection thresholds
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
-H 'Content-Type: application/json' \
-d '{"file_path":"uploads/piano.wav","onset_threshold":0.6,"frame_threshold":0.3,"minimum_note_length_ms":80,"output_path":"out/piano.mid"}'
# Write directly to a different staging path
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
-H 'Content-Type: application/json' \
-d '{"file_path":"recordings/bass.wav","output_path":"midi/bass_notes.mid"}'
# β {"path":"midi/bass_notes.mid","size":...,"engine":"basic-pitch","output_format":"mid"}
Optional params: onset_threshold (0β1, default 0.5), frame_threshold (0β1, default 0.3), minimum_note_length_ms (default 58), minimum_frequency / maximum_frequency (Hz, default unconstrained), multiple_pitch_bends (bool, default false), melodia_trick (bool, default true β helps with melodic content). Default engine: basic-pitch.
The MIDI file is piped straight into /v1/midi/inspect or /v1/midi/render β audio β MIDI β audio is a complete round-trip.
DeepFilterNet DF3 β deep learning noise suppression trained on speech. Better than bro
Content type
Image
Digest
sha256:ac4d6e91cβ¦
Size
1.7 GB
Last updated
3 days ago
docker pull psyb0t/audiolla