Sign inSign up

asgard/pdf-parser

By asgard

•Updated over 9 years ago

Spark app to parse scientific paper with grobid engine

Image
0

382

asgard/pdf-parser repository overview

⁠pdf-parser

Spark app to parse scientific paper with grobid engine

⁠build

docker build -t pdf-parser .

⁠usage

show usage:

docker run --rm pdf-parser spark-submit target/scala-2.10/grobid-spark-papers-assembly-1.6.0.jar

Usage: SparkPdfParser [options]

  --folderPath <value>
        path to folder containing pdf files to process
  --outputPath <value>
        path to output parquet file
  --numPartitions <value>
        Number of partitions of rdd to process
  --recursive
        Recursive flag to process inputFolder recursively of not. Default to false.
  --arxiv
        Flag to tell if files processed are ArXiv files in order to parse their id
  --test
        Flag to test the software, process only 2 pdf archive
  --yearFrom <value>
        Get file starting in yearFrom
  --yearTo <value>
        Get file until in yearTo

⁠run

docker run -p 5040:4040 --rm --name pdf-parser --net asgardapi_asgard_net pdf-parser spark-submit \
	--master spark://spark-master:7077 --driver-memory 3g --executor-memory 7g \
	--class pdfparser.SparkApp \
	/home/pdf-parser/target/scala-2.10/grobid-spark-papers-assembly-1.6.0.jar \
  	--folderPath s3a://asgard-data/test/arxiv/ArXiv_tar_backup \
  	--outputPath s3a://asgard-data/test/arxiv/arxiv3.parquet \
  	--numPartitions 4 \
  	--arxiv \
  	--test \
  	--yearFrom 2010 --yearTo 2017

Tag summary

Content type

Image

Digest

Size

2.6 GB

Last updated

over 9 years ago

docker pull asgard/pdf-parser