Text language identification using Wikipedia data
347
The aim of this project is to provide high-quality language detection over all the web's languages. The proxy for all web's languages is Wikipedia. Currently, we support 156 languages that have their Wikipedia entries.
The main function is text-langs that returns 2 values:
WILD> (text-langs "це тест")
((:UK . 0.5000003) (:RU . 0.4999998))
#(<це - UK:1.00> <тест - RU:1.00>)
$ cd wiki-lang-detect; sbcl --load run.lispdocker build -t wiki-lang-detect:latest .
docker run -it -p 5000:5000 wiki-lang-detect:latest
curl -X POST -H "Content-Type: application/json" -d "{'text': 'Несе Галя'}" http://localhost:5000/detect | jq '.'
Or you can use prebuilt Docker image maintained outside of this repository.
docker run -it -p 5000:5000 chaliy/wiki-lang-detect:latest
Content type
Image
Digest
Size
291.6 MB
Last updated
almost 10 years ago
docker pull chaliy/wiki-lang-detect