Sign inSign up

gpocali/music-identification

By gpocali

•Updated 3 months ago

A fully offline tool designed to create a report of music tracks in a video file

Image
0

2.8K

gpocali/music-identification repository overview

⁠Local Audio Recognizer

A fully offline, Dockerized acoustic fingerprinting tool designed to scan video files and identify which background tracks from your local music library are playing.

Built on top of ffmpeg and audfprint, this tool is specifically tuned to recognize music even when there is heavy voiceover or dialogue on top of it. It requires no external APIs, subscriptions, or internet access.

⁠Features

  • 100% Local Processing: Keeps your media and data entirely on your own hardware.
  • Voiceover Tolerant: Acoustic fingerprinting algorithm reliably detects background music beneath dialogue or ambient noise.
  • Recursive Scanning: Automatically crawls nested subdirectories to organize your reference tracks and video files however you like.
  • Chronological Reports: Generates automated, timestamped .txt reports mapping out exactly when each song begins and ends in your videos.
  • Persistent Database: Saves the audio fingerprint database locally to skip the hashing process on subsequent runs.
  • Runtime Configuration: Exposes a config.json file to adjust sensitivity and matching parameters on the fly without needing to rebuild the Docker image.

⁠Prerequisites

  • Docker installed on your host machine.
  • A directory of reference audio files (e.g., .mp3, .wav).
  • A directory of target video files (e.g., .mp4, .mkv).

⁠Quick Start (Docker Hub)

You do not need to download the source code to use this tool. You can run the pre-built image directly from Docker Hub.

  1. Create the data directories: The container expects a specific volume structure on your host machine. Create it by running:
    mkdir -p data/songs data/videos
    

2. **Add your media:**
* Place all reference music tracks in `data/songs/` (subdirectories are supported).
* Place all target video files in `data/videos/` (subdirectories are supported).


3. **Run the container:**
Execute the container, mapping your local `data` directory to the container's `/app/data` directory:
```bash
docker run --rm -v $(pwd)/data:/app/data gpocali/music-identification

On the first run, the script will:

  1. Generate a config.json file in your data directory with default high-sensitivity settings.
  2. Build the acoustic fingerprint database (fingerprints.pklz) from your reference songs.
  3. Scan your videos and output reports to data/reports/.

On subsequent runs, the tool skips the database generation phase and immediately begins scanning videos, drastically reducing runtime.

⁠Configuration (config.json)

The config.json file is automatically generated in your data directory. You can edit this file on your host machine to adjust the scanner's behavior for the next run.

{
    "density": "40",
    "min_count": "4",
    "shifts": "4",
    "rebuild_db": false
}

  • density: Increases the number of frequency peaks recorded per second during the database generation. 40 provides hyper-sensitivity for heavily obscured audio.
  • min_count: The minimum number of matching constellation peaks required to declare a successful match. Lowering this to 4 increases sensitivity but may introduce false positives if set any lower.
  • shifts: Forces the algorithm to check multiple sub-frame timing alignments. 4 ensures brief snippets aren't missed between standard frame windows.
  • rebuild_db: If you add new music to your songs directory, change this to true. On the next run, the container will rebuild fingerprints.pklz to include the new songs, and then automatically reset this flag to false.

⁠Expected Output

For every video scanned, a text file is generated in the data/reports/ directory. The report lists all identified tracks in chronological order based on their start time in the video.

Example Report (youtube_upload_final.mp4_report.txt):

Video File Evaluated: youtube_upload_final.mp4
==================================================
1. Start: 00:01:12 | End: 00:04:15 -> background_track_1.mp3
2. Start: 00:06:30 | End: 00:08:45 -> ambient_loop_2.wav
==================================================
Total number of songs used: 2

⁠Building from Source

If you wish to modify the code or build the image yourself:

  1. Clone the repository:
git clone [https://github.com/yourusername/local-audio-recognizer.git](https://github.com/yourusername/local-audio-recognizer.git)
cd local-audio-recognizer

  1. Build the Docker image locally:
docker build -t local-music-recognizer .

  1. Run your local build:
docker run --rm -v $(pwd)/data:/app/data local-music-recognizer

⁠Technologies Used

Tag summary

Content type

Image

Digest

sha256:a066b9cd2…

Size

469 MB

Last updated

3 months ago

docker pull gpocali/music-identification