Sign inSign up

lfoppiano/grobid

By lfoppiano

•Updated 4 days ago

Machine learning library for extracting, parsing and re-structuring raw PDF documents as XML.

Buildkit cache
Image
Machine learning & AI
Data science
11

1M+

lfoppiano/grobid repository overview

GROBID (or Grobid) means GeneRation Of BIbliographic Data.

GROBID is a machine learning library for extracting, parsing and re-structuring raw documents such as PDF into structured TEI-encoded documents with a particular focus on technical and scientific publications. First developments started in 2008 as a hobby. In 2011 the tool has been made available in open source. Work on GROBID has been steady as side project since the beginning and is expected to continue until at least 2020 :)

For more information https://github.com/kermitt2/grobid⁠

Full docker images (including deep learning models and embeddings can be pulled at grobid/grobid (https://hub.docker.com/r/grobid/grobid⁠)

Tag summary

Content type

Image

Digest

sha256:e10bcc570…

Size

488.3 MB

Last updated

4 days ago

docker pull lfoppiano/grobid:latest-crf