Sign inSign up

alumae/konele

By alumae

•Updated over 7 years ago

Image
0

318

alumae/konele repository overview

This image contains a kaldi-gstreamer-server (https://github.com/alumae/kaldi-gstreamer-server⁠) instance, preconfigured for Estonian general-purpose speech recognition. It contains the same models that are used by server that powers the Kõnele Android app and Dikteeri (https://bark.phon.ioc.ee/dikteeri/⁠) webapp. Use it if you don't want to use the public server for some reason.

Usage:

  • Install Docker CE
  • Install this image: docker pull alumae/konele
  • Start the container: docker run -p 4000:80 alumae/konele

Now, the speech recognition server is running on port 4000. To use it via its HTTP API, execute (in another terminal session):

curl -T sentence.ogg "http://localhost:4000/client/dynamic/recognize

To use it via the websocket-based API (via Python), you need to use client.py from https://github.com/alumae/kaldi-gstreamer-server/blob/master/kaldigstserver/client.py⁠ (it requires ws4py Python package):

python kaldigstserver/client.py  -u "ws://localhost:4000/client/ws/speech" sentence.ogg

You can also stream audio from microphone directly to the server:

arecord -f S16_LE -r 16000 | python kaldigstserver/client.py -u "ws://localhost:4000/client/ws/speech" -

By default, the server starts only one worker process. That means, only one audio stream can be processed at a time. If you want to increase the number of workers, use the num_workers variable when starting the container, e.g.:

docker run -p 4000:80  -e num_workers=2 alumae/konele

Tag summary

Content type

Image

Digest

Size

3.1 GB

Last updated

over 7 years ago

docker pull alumae/konele