A fully offline tool designed to create a report of music tracks in a video file
2.8K
A fully offline, Dockerized acoustic fingerprinting tool designed to scan video files and identify which background tracks from your local music library are playing.
Built on top of ffmpeg and audfprint, this tool is specifically tuned to recognize music even when there is heavy voiceover or dialogue on top of it. It requires no external APIs, subscriptions, or internet access.
.txt reports mapping out exactly when each song begins and ends in your videos.config.json file to adjust sensitivity and matching parameters on the fly without needing to rebuild the Docker image..mp3, .wav)..mp4, .mkv).You do not need to download the source code to use this tool. You can run the pre-built image directly from Docker Hub.
mkdir -p data/songs data/videos
2. **Add your media:**
* Place all reference music tracks in `data/songs/` (subdirectories are supported).
* Place all target video files in `data/videos/` (subdirectories are supported).
3. **Run the container:**
Execute the container, mapping your local `data` directory to the container's `/app/data` directory:
```bash
docker run --rm -v $(pwd)/data:/app/data gpocali/music-identification
On the first run, the script will:
config.json file in your data directory with default high-sensitivity settings.fingerprints.pklz) from your reference songs.data/reports/.On subsequent runs, the tool skips the database generation phase and immediately begins scanning videos, drastically reducing runtime.
config.json)The config.json file is automatically generated in your data directory. You can edit this file on your host machine to adjust the scanner's behavior for the next run.
{
"density": "40",
"min_count": "4",
"shifts": "4",
"rebuild_db": false
}
density: Increases the number of frequency peaks recorded per second during the database generation. 40 provides hyper-sensitivity for heavily obscured audio.min_count: The minimum number of matching constellation peaks required to declare a successful match. Lowering this to 4 increases sensitivity but may introduce false positives if set any lower.shifts: Forces the algorithm to check multiple sub-frame timing alignments. 4 ensures brief snippets aren't missed between standard frame windows.rebuild_db: If you add new music to your songs directory, change this to true. On the next run, the container will rebuild fingerprints.pklz to include the new songs, and then automatically reset this flag to false.For every video scanned, a text file is generated in the data/reports/ directory. The report lists all identified tracks in chronological order based on their start time in the video.
Example Report (youtube_upload_final.mp4_report.txt):
Video File Evaluated: youtube_upload_final.mp4
==================================================
1. Start: 00:01:12 | End: 00:04:15 -> background_track_1.mp3
2. Start: 00:06:30 | End: 00:08:45 -> ambient_loop_2.wav
==================================================
Total number of songs used: 2
If you wish to modify the code or build the image yourself:
git clone [https://github.com/yourusername/local-audio-recognizer.git](https://github.com/yourusername/local-audio-recognizer.git)
cd local-audio-recognizer
docker build -t local-music-recognizer .
docker run --rm -v $(pwd)/data:/app/data local-music-recognizer
Content type
Image
Digest
sha256:a066b9cd2…
Size
469 MB
Last updated
3 months ago
docker pull gpocali/music-identification