Welsh language speech recognition from fine-tuned wav2vec2 models on CPU devices.
397
e.g. in the same folder as your audio file and replace speech.wav with your audio file's filename
docker run --rm -it -v ${PWD}/:${PWD} techiaith/wav2vec2-inference-cpu python3 transcriber.py -w ${PWD}/speech.wav
Note the id of the video in the web address. E.g. https://www.youtube.com/watch?v=OpiwHxPPqRI for use in the command line:
docker run --rm -it -v ${PWD}/recordings/:/recordings techiaith/wav2vec2-inference-cpu yt.sh OpiwHxPPqRI
In the recordings sub-directory you will find a .srt, .TextGrid and a .wav audio file.
/home/<user>$ cd recordings/
/home/<user>/recordings$ ls -l
total 6708
-rw-r--r-- 1 root root 4737 Oct 5 17:01 OpiwHxPPqRI.srt
-rw-r--r-- 1 root root 17766 Oct 5 17:01 OpiwHxPPqRI.TextGrid
-rw-r--r-- 1 root root 6836776 Oct 5 17:00 OpiwHxPPqRI.wav
All models and images have been built from the content of https://github.com/techiaith/docker-wav2vec2-cy
Files for models are also available from the HuggingFace models hub. See https://huggingface.co/techiaith/wav2vec2-xlsr-ft-cy
Content type
Image
Digest
sha256:bfdbc6249…
Size
4.6 GB
Last updated
almost 4 years ago
docker pull techiaith/wav2vec2-inference-cpu