Spark app to parse scientific paper with grobid engine
382
Spark app to parse scientific paper with grobid engine
docker build -t pdf-parser .
show usage:
docker run --rm pdf-parser spark-submit target/scala-2.10/grobid-spark-papers-assembly-1.6.0.jar
Usage: SparkPdfParser [options]
--folderPath <value>
path to folder containing pdf files to process
--outputPath <value>
path to output parquet file
--numPartitions <value>
Number of partitions of rdd to process
--recursive
Recursive flag to process inputFolder recursively of not. Default to false.
--arxiv
Flag to tell if files processed are ArXiv files in order to parse their id
--test
Flag to test the software, process only 2 pdf archive
--yearFrom <value>
Get file starting in yearFrom
--yearTo <value>
Get file until in yearTo
docker run -p 5040:4040 --rm --name pdf-parser --net asgardapi_asgard_net pdf-parser spark-submit \
--master spark://spark-master:7077 --driver-memory 3g --executor-memory 7g \
--class pdfparser.SparkApp \
/home/pdf-parser/target/scala-2.10/grobid-spark-papers-assembly-1.6.0.jar \
--folderPath s3a://asgard-data/test/arxiv/ArXiv_tar_backup \
--outputPath s3a://asgard-data/test/arxiv/arxiv3.parquet \
--numPartitions 4 \
--arxiv \
--test \
--yearFrom 2010 --yearTo 2017
Content type
Image
Digest
Size
2.6 GB
Last updated
over 9 years ago
docker pull asgard/pdf-parser