Sign inSign up

rahulbot/cliff-clavin

By rahulbot

•Updated over 4 years ago

CLIFF-CLAVIN parses & disambiguates places mentioned in news articles.

Image
0

50K+

rahulbot/cliff-clavin repository overview

⁠CLIFF-CLAVIN

CLIFF-CLAVIN parses news articles and pulls out people, organizations and places mentioned. A number of tools do this, so why did we create CLIFF-CLAVIN? We've built on those tools to add disambiguation tailored to the ways news articles are written, and a concept of "focus" that tries to get at what place an article is really about (as opposed to all the places it mentions). We wrote CLIFF-CLAVIN to help drive our Media Cloud⁠ suite of tools, but are sharing it in hopes that others find it useful.

Note: CLAVIN, and by extension CLIFF, is very memory hungry due to the geonames index. To properly run, a minimum of 4GB of RAM is necessary. Any less and you'll experience errors.

⁠Running Via DockerHub

The quickest path to running CLIFF is to fetch the latest release from DockerHub:

docker pull rahulbot/cliff-clavin:latest
docker run -p 8080:8080 -m 8G -d rahulbot/cliff-clavin:latest

Then just hit a URL like this to see some JSON results:

http://localhost:8080/cliff-2.6.1/parse/text?q=This%20is%20some%20text%20about%20New%20York%20City,%20and%20maybe%20about%20Accra%20as%20well,%20and%20maybe%20Boston%20as%20well.

Notes:

⁠Acknowledgements

This is forked from John Beieler's cliff-docker⁠, which pulls heavily from Andy Halterman's CLIFF-up⁠ Vagrant box.

Tag summary

Content type

Image

Digest

Size

2.6 GB

Last updated

over 5 years ago

docker pull rahulbot/cliff-clavin