This image contains a kaldi-gstreamer-server (https://github.com/alumae/kaldi-gstreamer-server) instance, preconfigured for Estonian general-purpose speech recognition. It contains the same models that are used by server that powers the Kõnele Android app and Dikteeri (https://bark.phon.ioc.ee/dikteeri/) webapp. Use it if you don't want to use the public server for some reason.
Usage:
Now, the speech recognition server is running on port 4000. To use it via its HTTP API, execute (in another terminal session):
curl -T sentence.ogg "http://localhost:4000/client/dynamic/recognize
To use it via the websocket-based API (via Python), you need to use client.py from https://github.com/alumae/kaldi-gstreamer-server/blob/master/kaldigstserver/client.py (it requires ws4py Python package):
python kaldigstserver/client.py -u "ws://localhost:4000/client/ws/speech" sentence.ogg
You can also stream audio from microphone directly to the server:
arecord -f S16_LE -r 16000 | python kaldigstserver/client.py -u "ws://localhost:4000/client/ws/speech" -
By default, the server starts only one worker process. That means, only one audio stream can be processed at a time. If you want to increase the number of workers, use the num_workers variable when starting the container, e.g.:
docker run -p 4000:80 -e num_workers=2 alumae/konele
Content type
Image
Digest
Size
3.1 GB
Last updated
over 7 years ago
docker pull alumae/konele