Sign inSign up

footprintai/cockburncli

By footprintai

•Updated almost 2 years ago

Cockburn is a showcasing the prototyping of Micro Model Architecture with popular AI/ML solutions.

Machine learning & AI
2

99

footprintai/cockburncli repository overview

⁠cockburn

Cockburn is a monolithic repository showcasing the prototyping of Micro Model Architecture (MMA) technology with popular AI/ML solutions.

⁠Showcases

We have implemented a few showcases::

  • ASR (Aucustic speech recognition): This involves using AI/ML models to transcribe audio data into text. Notable solutions include:

    • Whisper⁠: An OpenAI model known for its high performance.
    • Whisper.cpp⁠, A C/C++ port of Whisper for faster computation.
    • MMA-variant of OpenAI Whisper.
    • MMA-variant of Whisper.cpp: We benchmarked this variant with 112 audio files (totaling 72 hours) in both MP3 and WAV format and achieved up to a 50% speed improvement with comparable performance. For details, see cli/whisper⁠.
  • Environment: RTX 3090, intel 13700k (10c), DDR4 3200 24GB.

    solutionmodelduration (s)
    mp3
    Whisper.cpplarge-v313,125 (s)
    Whisper.cpp + mmalarge-v36,982 (s)
    openai whisperlarge-v350,500 (s)
    openai whisper + mmalarge-v322,620 (s)
    faster whisperdistil-large-v28,762 (s)
    faster whisper + mmadistil-large-v25,880 (s)

⁠Quick start

⁠Flags

We support a list of flags listed below:

NAME:
   Whisper Speech-to-Text with Micro-model architecture - Speech-to-Text.

USAGE:
   Whisper Speech-to-Text with Micro-model architecture [global options] command [command options]

VERSION:
   1.0.9

AUTHOR:
    <[email protected]>

COMMANDS:
   help, h  Shows a list of commands or help for one command

GLOBAL OPTIONS:
   --model value                                    Model is the interface to a whisper model [$PLUGIN_MODEL, $INPUT_MODEL]
   --model-dir value                                ModelDir is a parent dir of the model (default: "/mnt/models") [$PLUGIN_MODEL_DIR, $INPUT_MODEL_DIR]
   --input-dir value                                input dir [$PLUGIN_INPUT_DIR, $INPUT_DIR]
   --input-regexp value                             input regexp for scanning, ex: .*.(mp3|mp4|wav) for scanning all mp3, mp4, and wav files (default: ".*.(mp3|mp4|wav)") [$PLUGIN_INPUT_REGEXP, $INPUT_REGEXP]
   --config value                                   config file [$PLUGIN_INPUT_CONFIG_FILE, $INPUT_CONFIG_FILE]
   --batch-input-enabled                            enable batch input mode (default: false) [$PLUGIN_BATCH_INPUT_ENABLED, $INPUT_BATCH_INPUT_ENABLED]
   --audio-path value                               audio path, comma delimited string [$PLUGIN_AUDIO_PATH, $INPUT_AUDIO_PATH]
   --output-folder value                            output folder [$PLUGIN_OUTPUT_FOLDER, $OUTPUT_FOLDER]
   --output-format value [ --output-format value ]  output format, support txt, srt, csv, json (default: "json") [$PLUGIN_OUTPUT_FORMAT, $OUTPUT_FORMAT]
   --language value                                 Set the language to use for speech recognition (default: "auto") [$PLUGIN_LANGUAGE, $INPUT_LANGUAGE]
   --threads value                                  Set number of threads to use (default: 10) [$PLUGIN_THREADS, $INPUT_THREADS]
   --debug                                          enable debug mode (default: false) [$PLUGIN_DEBUG, $INPUT_DEBUG]
   --translate                                      translate from source language to english (default: false) [$PLUGIN_TRANSLATE, $INPUT_TRANSLATE]
   --print-progress                                 print progress (default: true) [$PLUGIN_PRINT_PROGRESS, $INPUT_PRINT_PROGRESS]
   --print-segment                                  print segment (default: false) [$PLUGIN_PRINT_SEGMENT, $INPUT_PRINT_SEGMENT]
   --prompt value                                   initial prompt [$PLUGIN_PROMPT, $INPUT_PROMPT]
   --max-context value                              maximum number of text context tokens to store (default: 12) [$PLUGIN_MAX_CONTEXT, $INPUT_MAX_CONTEXT]
   --beam-size value                                beam size for beam search (default: 5) [$PLUGIN_BEAM_SIZE, $INPUT_BEAM_SIZE]
   --entropy-thold value                            entropy threshold for decoder fail (default: 2.4) [$PLUGIN_ENTROPY_THOLD, $INPUT_ENTROPY_THOLD]
   --whisper-engine value                           engine of whisper, choices: [whispercpp,fasterwhisper,openaiwhisper] (default: "whispercpp") [$PLUGIN_WHISPER_ENGINE, $WHISPER_ENGINE]
   --mma-parallelism value                          num of workers to run inference parallel (default: 2) [$PLUGIN_MMA_WORKERS, $MMA_PARALLELISM]
   --mma-concurrency value                          num of threads to run concurrently (default: 4) [$PLUGIN_MMA_CONCURRENCY, $MMA_CONCURRENCY]
   --port value                                     port for mma daemon (default: -1) [$PLUGIN_MMA_PORT, $MMA_PORT]
   --mma-batch-size value                           batch size for mma job (default: 4) (default: 4) [$PLUGIN_MMA_BATCH_SIZE, $MMA_BATCH_SIZE]
   --grpc-input-enabled                             enable grpc input (default: false) [$PLUGIN_GRPC_INPUT_ENABLED, $GRPC_INPUT_ENABLED]
   --grpc-input-port value                          port number for grpc service (default: -1) [$PLUGIN_GRPC_INPUT_PORT, $GRPC_INPUT_PORT]
   --vad-enabled value                              vad-enabled [$PLUGIN_VAD_ENABLED, $VAD_ENABLED]
   --vad-model-name value                           vad model name (default: "silero_vad") [$PLUGIN_VAD_MODEL_NAME, $VAD_MODEL_NAME]
   --wss-input-port value                           port number for websocket service (default: -1) [$PLUGIN_WSS_INPUT_PORT, $WSS_INPUT_PORT]
   --wss-input-enabled                              enable wss input (default: false) [$PLUGIN_WSS_INPUT_ENABLED, $INPUT_WSS_INPUT_ENABLED]
   --verbose                                        Verbose level (default: false) [$PLUGIN_LOG_VERBOSE, $INPUT_LOG_VERBOSE]
   --help, -h                                       show help
   --version, -v                                    print the version

COPYRIGHT:
   Copyright (c) 2024 Footprint AI
⁠Environment

our container image is build with the following configurations

Packageversion
ubuntu22.04
CUDA12.0.0
golang1.23.2
torch2.2.1
whisper.cpp1.7.1
faster-whisperv1.0.3
OpenAI Whisperv20240930
⁠Arguments
⁠Runtime Engine Selection
  • Run Faster Whisper, use --whisper-engine fasterwhisper
  • Run OpenAI Whisper, use --whisper-engine openaiwhisper
  • Run Whisper.cpp, use --whisper-engine whispercpp. (default)
⁠Inputs

We support three input methods:

  • Batch Inputs: Specify input files with full paths, separated by commas, for batch processing. For example, to load the large-v3 model and transcribe $fullpath1 and $fullpath2, outputting results in JSON format to the /output folder:

    --model /mnt/models/ggml-large-v3.bin \
    --whisper-engine whispercpp \
    --audio-path $fullpath1,$fullpath2 \
    --output-folder /output \
    --output-format json \
    
  • Config File: Use a JSON config file to list all inputs and outputs. For example, to specify inputs and outputs in a single JSON file and run them:

    [{
          "audio-path": "/data/wa3zOc_fjiI.mp3",
          "output-folder": "/output"
    }]
    

    and run them with --config configfile.json.

  • Scan Folder: Scan a folder for audio files (mp3, mp4, wav) using a regex pattern. For example, to scan the /data folder and process matching files:

    --input-dir /data \
    --input-regexp ".*.(mp3|mp4|wav)" \
    --output-folder /output
    
⁠Batch Inference
  • To run with a container, you can mount audio data (ex: /home/ubuneu/audiofiles and model folder (ex: home/ubuneu/modelfiles) to the container and run them directly. In this example, it granted container to run with 10(c) CPU and 24 G Ram.

    docker run --cpus=10 -m=24g --gpus all -it \
    	-v /home/ubuneu/modelfiles:/mnt/models \
    	-v /home/ubuntu/audiofiles:/data \
    	-v /home/ubuntu/output:/output \
    	footprintai/cockburncli:latest \
    		--input-dir=/data \
    		--input-regexp=".*.(mp3|mp4|wav)" \
    		--output-folder /output \
    		--whisper-engine whispercpp
    
  • To run with an OpenAI Whisper solution, use

    docker run --cpus=10 -m=24g  --gpus all -it \
    	-v /home/ubuneu/modelfiles:/mnt/models \
    	-v /home/ubuntu/audiofiles:/data \
    	-v /home/ubuntu/output:/output \
    	footprintai/cockburncli:latest \
    		--model large-v3 \
    		--input-dir /data \
    		--input-regexp ".*.(mp3|mp4|wav)" \
    		--output-folder /output \
    		--whisper-engine openaiwhisper
    
  • To run with an Faster Whisper solution, use

    docker run --cpus=10 -m=24g  --gpus all -it \
    	-v /home/ubuneu/modelfiles:/mnt/models \
    	-v /home/ubuntu/audiofiles:/data \
    	-v /home/ubuntu/output:/output \
    	footprintai/cockburncli:latest \
    		--model large-v3 \
    		--input-dir /data \
    		--input-regexp ".*.(mp3|mp4|wav)" \
    		--output-folder /output \
    		--whisper-engine fasterwhisper
    
⁠Websocket mode
  • To use websocket when you are about to to streaming inference, run the command with the following
docker run --gpus all -it -p 9000:9000 \
	-v /home/ubuneu/modelfiles:/mnt/models \
	footprintai/cockburncli:latest \
	--wss-input-enabled \
	--wss-input-port 9000 \
	--model base 
	--whisper-engine openaiwhisper 

this would open a port 9000 for websocket, and the frontend can reach with ws://localhost:9000

⁠Build by yourself
  • To build by your self, run make all from ./cli folder.

    cd cli/whisper && \
        make all
    
  • FFmpeg: this toolkit requires the command-line tool ffmpeg to be installed, it is available from mackage manager

    # on Ubuntu
    
    RUN apt-get update && \
      apt-get install -y --no-install-recommends curl ffmpeg \
      && rm -rf /var/lib/apt/lists/* /var/cache/apt/archives/*
    

⁠Roadmaps

  • Implements mma-vairant of llama.cpp

⁠FAQ

  • libcuda.so.1: cannot open shared object file. This means your host didn't have installed cuda driver, please install cuda and get a valid GPU to run it.

    root@72c2fa639ae9:/app# ./whisper -vvv
    ./whisper: error while loading shared libraries: libcuda.so.1: cannot open shared 	object file: No such file or directory
    

⁠Discussion

If you have any kind of feedback about this project feel free to use the Discussions section and open a new topic. If you have a question, make sure to check the Frequently asked questions and fire an issue. [email protected]⁠

Tag Summary

No tags have been pushed to this repository yet.