CLIFF-CLAVIN parses & disambiguates places mentioned in news articles.
50K+
CLIFF-CLAVIN parses news articles and pulls out people, organizations and places mentioned. A number of tools do this, so why did we create CLIFF-CLAVIN? We've built on those tools to add disambiguation tailored to the ways news articles are written, and a concept of "focus" that tries to get at what place an article is really about (as opposed to all the places it mentions). We wrote CLIFF-CLAVIN to help drive our Media Cloud suite of tools, but are sharing it in hopes that others find it useful.
Note: CLAVIN, and by extension CLIFF, is very memory hungry due to the geonames index. To properly run, a minimum of 4GB of RAM is necessary. Any less and you'll experience errors.
The quickest path to running CLIFF is to fetch the latest release from DockerHub:
docker pull rahulbot/cliff-clavin:latest
docker run -p 8080:8080 -m 8G -d rahulbot/cliff-clavin:latest
Then just hit a URL like this to see some JSON results:
http://localhost:8080/cliff-2.6.1/parse/text?q=This%20is%20some%20text%20about%20New%20York%20City,%20and%20maybe%20about%20Accra%20as%20well,%20and%20maybe%20Boston%20as%20well.
Notes:
This is forked from John Beieler's cliff-docker, which pulls heavily from Andy Halterman's CLIFF-up Vagrant box.
Content type
Image
Digest
Size
2.6 GB
Last updated
over 5 years ago
docker pull rahulbot/cliff-clavin