Cockburn is a showcasing the prototyping of Micro Model Architecture with popular AI/ML solutions.
99
Cockburn is a monolithic repository showcasing the prototyping of Micro Model Architecture (MMA) technology with popular AI/ML solutions.
We have implemented a few showcases::
ASR (Aucustic speech recognition): This involves using AI/ML models to transcribe audio data into text. Notable solutions include:
Environment: RTX 3090, intel 13700k (10c), DDR4 3200 24GB.
| solution | model | duration (s) |
|---|---|---|
| mp3 | ||
| Whisper.cpp | large-v3 | 13,125 (s) |
| Whisper.cpp + mma | large-v3 | 6,982 (s) |
| openai whisper | large-v3 | 50,500 (s) |
| openai whisper + mma | large-v3 | 22,620 (s) |
| faster whisper | distil-large-v2 | 8,762 (s) |
| faster whisper + mma | distil-large-v2 | 5,880 (s) |
We support a list of flags listed below:
NAME:
Whisper Speech-to-Text with Micro-model architecture - Speech-to-Text.
USAGE:
Whisper Speech-to-Text with Micro-model architecture [global options] command [command options]
VERSION:
1.0.9
AUTHOR:
<[email protected]>
COMMANDS:
help, h Shows a list of commands or help for one command
GLOBAL OPTIONS:
--model value Model is the interface to a whisper model [$PLUGIN_MODEL, $INPUT_MODEL]
--model-dir value ModelDir is a parent dir of the model (default: "/mnt/models") [$PLUGIN_MODEL_DIR, $INPUT_MODEL_DIR]
--input-dir value input dir [$PLUGIN_INPUT_DIR, $INPUT_DIR]
--input-regexp value input regexp for scanning, ex: .*.(mp3|mp4|wav) for scanning all mp3, mp4, and wav files (default: ".*.(mp3|mp4|wav)") [$PLUGIN_INPUT_REGEXP, $INPUT_REGEXP]
--config value config file [$PLUGIN_INPUT_CONFIG_FILE, $INPUT_CONFIG_FILE]
--batch-input-enabled enable batch input mode (default: false) [$PLUGIN_BATCH_INPUT_ENABLED, $INPUT_BATCH_INPUT_ENABLED]
--audio-path value audio path, comma delimited string [$PLUGIN_AUDIO_PATH, $INPUT_AUDIO_PATH]
--output-folder value output folder [$PLUGIN_OUTPUT_FOLDER, $OUTPUT_FOLDER]
--output-format value [ --output-format value ] output format, support txt, srt, csv, json (default: "json") [$PLUGIN_OUTPUT_FORMAT, $OUTPUT_FORMAT]
--language value Set the language to use for speech recognition (default: "auto") [$PLUGIN_LANGUAGE, $INPUT_LANGUAGE]
--threads value Set number of threads to use (default: 10) [$PLUGIN_THREADS, $INPUT_THREADS]
--debug enable debug mode (default: false) [$PLUGIN_DEBUG, $INPUT_DEBUG]
--translate translate from source language to english (default: false) [$PLUGIN_TRANSLATE, $INPUT_TRANSLATE]
--print-progress print progress (default: true) [$PLUGIN_PRINT_PROGRESS, $INPUT_PRINT_PROGRESS]
--print-segment print segment (default: false) [$PLUGIN_PRINT_SEGMENT, $INPUT_PRINT_SEGMENT]
--prompt value initial prompt [$PLUGIN_PROMPT, $INPUT_PROMPT]
--max-context value maximum number of text context tokens to store (default: 12) [$PLUGIN_MAX_CONTEXT, $INPUT_MAX_CONTEXT]
--beam-size value beam size for beam search (default: 5) [$PLUGIN_BEAM_SIZE, $INPUT_BEAM_SIZE]
--entropy-thold value entropy threshold for decoder fail (default: 2.4) [$PLUGIN_ENTROPY_THOLD, $INPUT_ENTROPY_THOLD]
--whisper-engine value engine of whisper, choices: [whispercpp,fasterwhisper,openaiwhisper] (default: "whispercpp") [$PLUGIN_WHISPER_ENGINE, $WHISPER_ENGINE]
--mma-parallelism value num of workers to run inference parallel (default: 2) [$PLUGIN_MMA_WORKERS, $MMA_PARALLELISM]
--mma-concurrency value num of threads to run concurrently (default: 4) [$PLUGIN_MMA_CONCURRENCY, $MMA_CONCURRENCY]
--port value port for mma daemon (default: -1) [$PLUGIN_MMA_PORT, $MMA_PORT]
--mma-batch-size value batch size for mma job (default: 4) (default: 4) [$PLUGIN_MMA_BATCH_SIZE, $MMA_BATCH_SIZE]
--grpc-input-enabled enable grpc input (default: false) [$PLUGIN_GRPC_INPUT_ENABLED, $GRPC_INPUT_ENABLED]
--grpc-input-port value port number for grpc service (default: -1) [$PLUGIN_GRPC_INPUT_PORT, $GRPC_INPUT_PORT]
--vad-enabled value vad-enabled [$PLUGIN_VAD_ENABLED, $VAD_ENABLED]
--vad-model-name value vad model name (default: "silero_vad") [$PLUGIN_VAD_MODEL_NAME, $VAD_MODEL_NAME]
--wss-input-port value port number for websocket service (default: -1) [$PLUGIN_WSS_INPUT_PORT, $WSS_INPUT_PORT]
--wss-input-enabled enable wss input (default: false) [$PLUGIN_WSS_INPUT_ENABLED, $INPUT_WSS_INPUT_ENABLED]
--verbose Verbose level (default: false) [$PLUGIN_LOG_VERBOSE, $INPUT_LOG_VERBOSE]
--help, -h show help
--version, -v print the version
COPYRIGHT:
Copyright (c) 2024 Footprint AI
our container image is build with the following configurations
| Package | version |
|---|---|
| ubuntu | 22.04 |
| CUDA | 12.0.0 |
| golang | 1.23.2 |
| torch | 2.2.1 |
| whisper.cpp | 1.7.1 |
| faster-whisper | v1.0.3 |
| OpenAI Whisper | v20240930 |
--whisper-engine fasterwhisper--whisper-engine openaiwhisper--whisper-engine whispercpp. (default)We support three input methods:
Batch Inputs: Specify input files with full paths, separated by commas, for batch processing. For example, to load the large-v3 model and transcribe $fullpath1 and $fullpath2, outputting results in JSON format to the /output folder:
--model /mnt/models/ggml-large-v3.bin \
--whisper-engine whispercpp \
--audio-path $fullpath1,$fullpath2 \
--output-folder /output \
--output-format json \
Config File: Use a JSON config file to list all inputs and outputs. For example, to specify inputs and outputs in a single JSON file and run them:
[{
"audio-path": "/data/wa3zOc_fjiI.mp3",
"output-folder": "/output"
}]
and run them with --config configfile.json.
Scan Folder: Scan a folder for audio files (mp3, mp4, wav) using a regex pattern. For example, to scan the /data folder and process matching files:
--input-dir /data \
--input-regexp ".*.(mp3|mp4|wav)" \
--output-folder /output
To run with a container, you can mount audio data (ex: /home/ubuneu/audiofiles and model folder (ex: home/ubuneu/modelfiles) to the container and run them directly.
In this example, it granted container to run with 10(c) CPU and 24 G Ram.
docker run --cpus=10 -m=24g --gpus all -it \
-v /home/ubuneu/modelfiles:/mnt/models \
-v /home/ubuntu/audiofiles:/data \
-v /home/ubuntu/output:/output \
footprintai/cockburncli:latest \
--input-dir=/data \
--input-regexp=".*.(mp3|mp4|wav)" \
--output-folder /output \
--whisper-engine whispercpp
To run with an OpenAI Whisper solution, use
docker run --cpus=10 -m=24g --gpus all -it \
-v /home/ubuneu/modelfiles:/mnt/models \
-v /home/ubuntu/audiofiles:/data \
-v /home/ubuntu/output:/output \
footprintai/cockburncli:latest \
--model large-v3 \
--input-dir /data \
--input-regexp ".*.(mp3|mp4|wav)" \
--output-folder /output \
--whisper-engine openaiwhisper
To run with an Faster Whisper solution, use
docker run --cpus=10 -m=24g --gpus all -it \
-v /home/ubuneu/modelfiles:/mnt/models \
-v /home/ubuntu/audiofiles:/data \
-v /home/ubuntu/output:/output \
footprintai/cockburncli:latest \
--model large-v3 \
--input-dir /data \
--input-regexp ".*.(mp3|mp4|wav)" \
--output-folder /output \
--whisper-engine fasterwhisper
docker run --gpus all -it -p 9000:9000 \
-v /home/ubuneu/modelfiles:/mnt/models \
footprintai/cockburncli:latest \
--wss-input-enabled \
--wss-input-port 9000 \
--model base
--whisper-engine openaiwhisper
this would open a port 9000 for websocket, and the frontend can reach with ws://localhost:9000
To build by your self, run make all from ./cli folder.
cd cli/whisper && \
make all
FFmpeg: this toolkit requires the command-line tool ffmpeg to be installed, it is available from mackage manager
# on Ubuntu
RUN apt-get update && \
apt-get install -y --no-install-recommends curl ffmpeg \
&& rm -rf /var/lib/apt/lists/* /var/cache/apt/archives/*
libcuda.so.1: cannot open shared object file. This means your host didn't have installed cuda driver, please install cuda and get a valid GPU to run it.
root@72c2fa639ae9:/app# ./whisper -vvv
./whisper: error while loading shared libraries: libcuda.so.1: cannot open shared object file: No such file or directory
If you have any kind of feedback about this project feel free to use the Discussions section and open a new topic. If you have a question, make sure to check the Frequently asked questions and fire an issue. [email protected]
No tags have been pushed to this repository yet.