Sign inSign up

psyb0t/audiolla

By psyb0t

β€’Updated 3 days ago

Image
0

9.5K

psyb0t/audiolla repository overview

⁠audiolla

CI version license Docker Pulls

Thirty audio engines. One port. Zero cloud. Fire-and-forget async jobs. Webhooks.

You needed Demucs for stems. Then librosa for BPM and key. Then basic-pitch for MIDI transcription. Then pyannote for speaker diarization. Then DeepFilterNet for speech enhancement. Then you spent three days debugging Python version conflicts and now you hate everything.

audiolla is what happens when you stop doing that.

Every audio processing tool worth using β€” wrapped in one HTTP API, running in one Docker container. POST a file. Get audio, JSON, or MIDI back. Drive it from curl, shell scripts, Python notebooks, Makefiles, or point an LLM agent at the MCP endpoint and let it rip.

No account. No subscription. No per-minute billing. No vendor lock-in. docker run and you're done.


⁠What's in the box

πŸŽ›οΈ Stem separationDemucs β€” htdemucs, fine-tuned, 6-stem, MDX variants
🎚️ MasteringReference mastering (matchering) + custom pedalboard chains
πŸ“Š AnalysisBPM Β· key Β· LUFS Β· beats Β· onsets Β· melody Β· structural segments
🎹 Chords + keyChord detection + Krumhansl-Schmuckler key estimation
🎡 Audio β†’ MIDIPolyphonic transcription via Spotify's basic-pitch (ONNX, no TF)
🧹 RestorationDe-reverb · de-echo · de-noise via UVR BS-Roformer + MelBand Roformer
πŸ—£οΈ SpeechEnhancement (DeepFilterNet) Β· VAD (silero-vad) Β· diarization (pyannote)
πŸ–ΌοΈ VisualsSpectrogram + waveform PNGs + 8-mode animated MP4/WebM
πŸ” FingerprintChromaprint acoustic fingerprinting (AcoustID-compatible)
βœ‚οΈ SilenceDetect gaps Β· trim edges Β· strip all silence
🎼 MIDI pipelineCompose from JSON · inspect · transform · render via fluidsynth
🎸 Effects23-effect pedalboard chain β€” Compressor, Reverb, PitchShift, filters…
πŸ”§ TransformsSox DSP β€” pitch, tempo, EQ, reverb, gain
πŸ“’ LoudnessMeasure LUFS Β· normalize to target
πŸ₯ HPSSHarmonic/percussive source separation via librosa median filter
πŸ”‡ Noise reductionSpectral noise reduction via noisereduce β€” stationary + adaptive modes
⏩ Time-stretchIndependent tempo factor + pitch shift via librosa phase vocoder
🏷️ Audio taggingTop-K AudioSet class labels via Audio Spectrogram Transformer
πŸ”— Audio embeddings512-dim semantic embeddings via LAION CLAP + optional text similarity
🏷️ Zero-shot classifyCLAP cosine similarity against any free-form text labels β€” genres, moods, instruments
πŸ“‹ Audio infoffprobe metadata β€” duration, sample rate, channels, codec, bit depth
βœ‚οΈ TrimCut a clip by start/end seconds β€” any format in, any format out
🎚️ MixCombine N staged tracks with per-track gain_db β€” pure ffmpeg, no model
πŸ”— ConcatStitch N audio files end-to-end in order
⏩ SpeedChange playback speed without pitch shift (0.1Γ— – 10Γ—) via ffmpeg atempo
πŸ”„ ConvertRe-encode: format, sample rate, channel count in one call
πŸ” SimilarCosine similarity between two audio files via CLAP embeddings
🎹 MIDI quantizeSnap MIDI note timings to a rhythmic grid (16th, 8th, quarter…)
πŸŒ… FadeFade-in and/or fade-out with 13 curve shapes
βͺ ReverseFlip audio backwards
πŸ” LoopRepeat audio N times
🎯 BPM matchAuto-detect BPM then stretch to a target β€” no manual math
πŸ“ˆ Loudness curveRMS envelope over time β€” time-stamped dB values for gain automation
🎀 Pitch correctAuto-tune toward nearest chromatic semitone β€” configurable strength
πŸ”§ RepairDeclip + dehum β€” fix clipped peaks and remove power-line hum
πŸ” Loop pointFind best seamless loop boundary β€” score, bar count, candidates list
πŸ₯ Drum machineStep-sequencer spec β†’ GM drum MIDI β€” 16-step pattern, swing, tempo
🎼 Chords to MIDIChord progression β†’ MIDI file β€” root+3rd+5th voicings per segment
↔️ Stereo widthWiden or collapse the stereo image via M/S processing
βœ‚οΈ SplitSplit into N equal parts or on silence β€” returns ZIP of segments
πŸ”Š PanPosition audio in the stereo field (-1 left β†’ 0 center β†’ 1 right)
🎚️ EQParametric EQ β€” JSON array of freq/gain_db/width_hz bands
🎡 Key matchDetect source key then pitch-shift to a target key
πŸŽ™οΈ Sidechain duckDuck music when a trigger track (voice) is loud
🏷️ MetadataRead and write ID3/Vorbis/FLAC/WAV audio tags via mutagen
πŸ”΄ Clip detectDetect digital clipping β€” count, ratio, peak dBFS
↔️ Mid/SideEncode L/R β†’ Mid+Side or decode Mid+Side β†’ L/R
βœ‚οΈ Beat sliceSlice audio at detected beat positions β€” returns ZIP of segments
🏟️ Conv reverbConvolution reverb via impulse response β€” wet_mix control
πŸ₯ Transient shaperAttack/sustain dual-compressor β€” punch up drums, cut room tail
🎚️ Multiband compressN-band compressor with zero-phase LR4 crossovers β€” mastering-grade dynamics
πŸŽ›οΈ DJ prepOne call: BPM + key + Camelot wheel position + integrated LUFS
πŸ“¦ BatchRun trim/convert/fade/reverse/speed/eq on staged files in sequence
🧩 Presets + pipelineCurated YAML workflows (master-for-spotify, podcast-cleanup, …) + ad-hoc op chaining server-side
πŸ—‚οΈ CatalogGET /v1/catalog β€” machine-readable endpoint list grouped by category for discovery
⚑ Async jobsEvery endpoint supports async_job=true β€” fire-and-forget + webhook callbacks

⁠Table of Contents


⁠Run it

# no GPU
docker run --rm -it \
  -v $HOME/.audiolla-data:/data \
  -p 8000:8000 \
  psyb0t/audiolla:latest

# GPU
docker run --rm -it --gpus all \
  -v $HOME/.audiolla-data:/data \
  -e AUDIOLLA_DEVICE=cuda \
  -p 8000:8000 \
  psyb0t/audiolla:latest-cuda

Demucs weights prefetch at container startup (for whichever variants are enabled) and cache in /data/torch_cache/. First boot downloads them; same -v mount next time and they're already there. Other engines (matchering, pedalboard, librosa, sox, fx, midi) have no weights β€” they're ready as soon as /healthz is green.


⁠Migration from v0.23.x β†’ v1.0.0

v1.0.0 is a breaking API release. Every existing client breaks. The new shape:

  • Every audio endpoint takes a JSON body (no more multipart/form-data except at /v1/files)
  • Input is file_path (FILES_DIR-relative) xor file_url (server-side fetch). Pre-stage the file via PUT /v1/files/{path} first.
  • Output requires output_path xor output_url. No more raw audio bytes in responses.
  • Async path: async_job=true auto-stages to jobs/{id}.{ext} if neither output is given.
  • MCP audio-producing tools dropped audio_base64 (and midi_base64 / image_base64 / video_base64). Same output_path xor output_url requirement.
  • openapi.yaml is now the contract β€” Pydantic models regenerate from it via make generate. Never hand-edit src/audiolla/schema/_generated.py.
- curl -X POST http://localhost:8000/v1/audio/normalize \
-     -F "[email protected]" -F "target_lufs=-14" -o normalized.wav

+ # 1) stage the file (multipart only lives here now)
+ curl -X PUT --data-binary @track.wav \
+     -H 'Content-Type: application/octet-stream' \
+     http://localhost:8000/v1/files/uploads/track.wav

+ # 2) process via JSON body β€” response is JSON, not bytes
+ curl -X POST http://localhost:8000/v1/audio/normalize \
+     -H 'Content-Type: application/json' \
+     -d '{"file_path":"uploads/track.wav","target_lufs":-14,"output_path":"out/normalized.wav"}'

+ # 3) retrieve the result
+ curl -o normalized.wav http://localhost:8000/v1/files/out/normalized.wav

Why? See the v1.0.0 CHANGELOG entry⁠ for the full rationale.

⁠Quick start

Once the container is up, this is a complete audio pipeline in six commands (every audio endpoint is JSON-body now; stage your input file at /v1/files/... first):

# stage your input file
curl -X PUT --data-binary @song.wav \
  -H 'Content-Type: application/octet-stream' \
  http://localhost:8000/v1/files/uploads/song.wav

# rip the vocals out of a track
curl -X POST http://localhost:8000/v1/audio/separate \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/song.wav","engine":"htdemucs","stems":["vocals"],"output_path":"out/vocals.wav"}'
# β†’ {"path":"out/vocals.wav","size":...,"output_format":"wav"}
curl -o vocals.wav http://localhost:8000/v1/files/out/vocals.wav

# what key is it in? what are the chords?
curl -X POST http://localhost:8000/v1/audio/chords \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/song.wav"}'
# β†’ {"key":"F# minor","key_confidence":0.91,"chords":[{"chord":"F#m","start_sec":0.0,...},...]}

# transcribe that vocal melody to MIDI
curl -X PUT --data-binary @out/vocals.wav -H 'Content-Type: application/octet-stream' \
  http://localhost:8000/v1/files/uploads/vocals.wav  # only if not already staged
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/vocals.wav","output_path":"out/melody.mid"}'

# render the MIDI back to audio through a SoundFont
curl -X POST http://localhost:8000/v1/midi/render \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"out/melody.mid","output_path":"out/rendered.wav"}'

# strip background noise from a voice recording
curl -X POST http://localhost:8000/v1/audio/noise-reduce/uvr-denoise \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/interview.wav","output_path":"out/clean.wav"}'

# who's speaking and when?
curl -X POST http://localhost:8000/v1/audio/diarize/pyannote \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/interview.wav"}'
# β†’ {"num_speakers":2,"segments":[{"speaker":"SPEAKER_00","start_sec":0.5,"end_sec":8.2},...]}

Audio in. MIDI out. Chords detected. Speakers identified. De-noised. Re-synthesized. No Python environment to set up. No API keys. No account. Just HTTP.


⁠What it can do

Output defaults to wav. Add "output_format":"mp3" to the JSON body to get mp3 instead (flac, opus, aac, pcm also work).

Every audio endpoint takes an application/json body. The only place multipart still lives is PUT /v1/files/{path} (raw bytes for staging an input file).

Input β€” every audio endpoint requires exactly one of:

  • file_path β€” path inside the /v1/files staging area (stage with PUT /v1/files/{path} first)
  • file_url β€” remote URL the server fetches (disabled by default β€” see Remote URLs⁠)

Output β€” audio-producing endpoints require exactly one of:

  • output_path β€” server writes to /v1/files/<path>, returns JSON {"path":..., "size":..., ...}
  • output_url β€” server PUTs to a presigned URL, returns JSON {"url":..., "size":..., ...}

Analysis-only endpoints (those that return JSON data, e.g. /v1/audio/analyze, /v1/audio/loudness, /v1/audio/info) don't need output_path / output_url β€” the response is the result.

⁠Split stems
# stage input
curl -X PUT --data-binary @track.wav \
  -H 'Content-Type: application/octet-stream' \
  http://localhost:8000/v1/files/uploads/track.wav

# vocals only
curl -X POST http://localhost:8000/v1/audio/separate \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","engine":"htdemucs","stems":["vocals"],"output_path":"out/vocals.wav"}'
curl -o vocals.wav http://localhost:8000/v1/files/out/vocals.wav

# all 4 stems as a ZIP
curl -X POST http://localhost:8000/v1/audio/separate \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","engine":"htdemucs","output_path":"out/stems.zip"}'
curl -o stems.zip http://localhost:8000/v1/files/out/stems.zip
⁠Master
# stage track + reference
curl -X PUT --data-binary @track.wav -H 'Content-Type: application/octet-stream' \
  http://localhost:8000/v1/files/uploads/track.wav
curl -X PUT --data-binary @ref.wav -H 'Content-Type: application/octet-stream' \
  http://localhost:8000/v1/files/uploads/ref.wav

# match EQ + loudness to a reference track
curl -X POST http://localhost:8000/v1/audio/master \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","mode":"reference","reference_path":"uploads/ref.wav","output_path":"out/mastered.wav"}'
curl -o mastered.wav http://localhost:8000/v1/files/out/mastered.wav

# run a built-in pedalboard chain (presets: transparent, loud)
curl -X POST http://localhost:8000/v1/audio/master \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","mode":"chain","preset":"loud","output_path":"out/mastered.wav"}'
curl -o mastered.wav http://localhost:8000/v1/files/out/mastered.wav
⁠Analyze
# returns JSON. features: bpm, key, loudness, duration,
# spectral_centroid, rms, zcr. Omit features to get them all.
curl -X POST http://localhost:8000/v1/audio/analyze \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","features":["bpm","key","loudness"]}'
⁠Beats, onsets, melody, segments
# beat grid β€” returns bpm + beat timestamps
curl -X POST http://localhost:8000/v1/audio/beats \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav"}'

# onset timestamps β€” note attacks, transients
curl -X POST http://localhost:8000/v1/audio/onsets \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav"}'

# dominant melody contour β€” pitch in Hz per frame
curl -X POST http://localhost:8000/v1/audio/melody \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav"}'

# structural segmentation β€” labels recurring sections A, B, C...
curl -X POST http://localhost:8000/v1/audio/segments \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","num_segments":6}'

Beat detection also generates a click-track file when click_track=true (set output_path to receive it) β€” handy for aligning a mix to a grid. Pass start_bpm=140 to seed the tracker when you already know the rough tempo (faster, more accurate). Melody can be exported as a single-track MIDI file via as_midi=true + output_path.

⁠Silence detection and trimming
# find silent gaps in a recording (no trim_mode β†’ JSON only)
curl -X POST http://localhost:8000/v1/audio/silence \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","threshold_db":-30,"min_duration_sec":1.0}'

# trim all silence and stage the result
curl -X POST http://localhost:8000/v1/audio/silence \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","threshold_db":-30,"min_duration_sec":0.5,"trim_mode":"all","output_path":"out/trimmed.wav"}'
curl -o trimmed.wav http://localhost:8000/v1/files/out/trimmed.wav

# trim only leading/trailing silence
curl -X POST http://localhost:8000/v1/audio/silence \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","threshold_db":-40,"min_duration_sec":0.3,"trim_mode":"edges","output_path":"processed/trimmed.wav"}'

trim_mode=edges β€” chop leading + trailing silence only. trim_mode=all β€” remove every detected gap (compress a talk recording, tighten a loop). Without trim_mode, the response is JSON only: silent_ranges, non_silent_ranges, duration β€” and output_path / output_url is not required.

⁠Visualize (spectrogram, waveform, video)

Visual output splits into two sub-namespaces by output type:

# Static PNG spectrogram (color + scale params)
curl -X POST http://localhost:8000/v1/audio/visualize/image/spectrogram \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","width":1280,"height":720,"output_path":"out/spec.png"}'
curl -o spec.png http://localhost:8000/v1/files/out/spec.png

# Static PNG waveform (color param)
curl -X POST http://localhost:8000/v1/audio/visualize/image/waveform \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","width":1280,"height":240,"output_path":"out/wave.png"}'
curl -o wave.png http://localhost:8000/v1/files/out/wave.png

# Animated MP4 spectrum analyser (fps + container params)
curl -X POST http://localhost:8000/v1/audio/visualize/video/spectrum \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","width":1280,"height":720,"fps":30,"container":"mp4","output_path":"out/viz.mp4"}'
curl -o viz.mp4 http://localhost:8000/v1/files/out/viz.mp4

/image/spectrogram: produces a PNG (staged via output_path or PUT to output_url). Params: width, height, color (default intensity), scale (log/lin).

/image/waveform: produces a PNG. Params: width, height, color (default lime).

/video/{mode}: spectrum (scrolling FFT), waves (oscilloscope), cqt (constant-Q transform), freqs (bar-graph analyzer), volume (VU meter), vectorscope (stereo X/Y scope), phasemeter, histogram. Params: width, height, fps, container (mp4 default, webm).

⁠Acoustic fingerprint
# Chromaprint fingerprint β€” identifies a recording regardless of encoding
curl -X POST http://localhost:8000/v1/audio/fingerprint \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav"}'
# β†’ {"duration": 215.34, "fingerprint": "AQADtEqRRIuQ..."}

# include the raw integer array (for custom similarity scoring)
curl -X POST http://localhost:8000/v1/audio/fingerprint \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","return_raw":true}'

The base64 fingerprint string is compatible with the AcoustID⁠ lookup service.

⁠De-reverb, de-echo, de-noise

AI audio restoration via UVR ecosystem models β€” BS-Roformer and MelBand Roformer. All three are unified under POST /v1/audio/restore/{engine}.

# Remove room reverb (BS-Roformer, SDR 19+)
curl -X POST http://localhost:8000/v1/audio/restore/uvr-dereverb \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","output_path":"out/dry.wav"}'

# Remove echo β€” normal mode
curl -X POST http://localhost:8000/v1/audio/restore/uvr-deecho \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","output_path":"out/noecho.wav"}'

# Remove echo β€” aggressive mode (same engine, harder suppression)
curl -X POST http://localhost:8000/v1/audio/restore/uvr-deecho \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","aggressive":true,"output_path":"out/noecho.wav"}'

# Remove broadband background noise β€” ML (MelBand Roformer, SDR 28)
curl -X POST http://localhost:8000/v1/audio/restore/uvr-denoise \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/track.wav","output_path":"out/clean.wav"}'

All support output_format, output_path, output_url. For DSP-based noise reduction (no GPU) use noise-reduce/noise-reduce.

UVR engines also work through /v1/audio/separate β€” uvr-vocal-bsr (BS-Roformer, SDR 13) and uvr-karaoke return vocal + instrumental stems like Demucs but often with higher quality.

⁠Audio-to-MIDI transcription

Polyphonic audio-to-MIDI via Spotify's basic-pitch (ONNX backend, no TensorFlow). Play guitar, hum a melody, record a piano riff β€” get a MIDI file back with all the notes.

# Any audio β†’ MIDI file (staged)
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/guitar_riff.wav","output_path":"out/riff.mid"}'
curl -o riff.mid http://localhost:8000/v1/files/out/riff.mid

# Tune the detection thresholds
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"uploads/piano.wav","onset_threshold":0.6,"frame_threshold":0.3,"minimum_note_length_ms":80,"output_path":"out/piano.mid"}'

# Write directly to a different staging path
curl -X POST http://localhost:8000/v1/audio/to_midi/basic-pitch \
  -H 'Content-Type: application/json' \
  -d '{"file_path":"recordings/bass.wav","output_path":"midi/bass_notes.mid"}'
# β†’ {"path":"midi/bass_notes.mid","size":...,"engine":"basic-pitch","output_format":"mid"}

Optional params: onset_threshold (0–1, default 0.5), frame_threshold (0–1, default 0.3), minimum_note_length_ms (default 58), minimum_frequency / maximum_frequency (Hz, default unconstrained), multiple_pitch_bends (bool, default false), melodia_trick (bool, default true β€” helps with melodic content). Default engine: basic-pitch.

The MIDI file is piped straight into /v1/midi/inspect or /v1/midi/render β€” audio β†’ MIDI β†’ audio is a complete round-trip.

⁠Neural speech and vocal enhancement

DeepFilterNet DF3 β€” deep learning noise suppression trained on speech. Better than bro

Tag summary

Content type

Image

Digest

sha256:ac4d6e91c…

Size

1.7 GB

Last updated

3 days ago

docker pull psyb0t/audiolla