Sign inSign up

techiaith/wav2vec2-inference-cpu

By techiaith

•Updated almost 4 years ago

Welsh language speech recognition from fine-tuned wav2vec2 models on CPU devices.

Image
0

397

techiaith/wav2vec2-inference-cpu repository overview

⁠To transcribe a single file:

e.g. in the same folder as your audio file and replace speech.wav with your audio file's filename

docker run --rm -it -v ${PWD}/:${PWD} techiaith/wav2vec2-inference-cpu python3 transcriber.py -w ${PWD}/speech.wav

⁠To transcribe a YouTube video

Note the id of the video in the web address. E.g. https://www.youtube.com/watch?v=OpiwHxPPqRI⁠ for use in the command line:

docker run --rm -it -v ${PWD}/recordings/:/recordings techiaith/wav2vec2-inference-cpu yt.sh OpiwHxPPqRI

In the recordings sub-directory you will find a .srt, .TextGrid and a .wav audio file.

/home/<user>$ cd recordings/
/home/<user>/recordings$ ls -l
total 6708
-rw-r--r-- 1 root root    4737 Oct  5 17:01 OpiwHxPPqRI.srt
-rw-r--r-- 1 root root   17766 Oct  5 17:01 OpiwHxPPqRI.TextGrid
-rw-r--r-- 1 root root 6836776 Oct  5 17:00 OpiwHxPPqRI.wav

⁠Model and Source files

All models and images have been built from the content of https://github.com/techiaith/docker-wav2vec2-cy⁠

Files for models are also available from the HuggingFace models hub. See https://huggingface.co/techiaith/wav2vec2-xlsr-ft-cy⁠

Tag summary

Content type

Image

Digest

sha256:bfdbc6249…

Size

4.6 GB

Last updated

almost 4 years ago

docker pull techiaith/wav2vec2-inference-cpu