Project build around : https://github.com/idio/json-wikipedia
This docker image converts uncompressed XML dumps to JSON-per-line file. It uses Spark to parallelise the process, so the image explicitly sets minimum RAM to 20G - it won't start if this amount of memory is unavailable.
input path to uncompressed wikipedia xml dumpoutput path where the jsonwikipedia will be generateddocker run -v /tmp:/tmp -v /mnt:/mnt -i -t idio/jsonwikipedia -input /mnt/enwiki-20150602-pages-articles.xml -output /mnt/enwiki.json -lang en -action export-parallel
Your /tmp folder should have enough disk space.
Content type
Image
Digest
sha256:142b66ecb…
Size
1.2 GB
Last updated
about 11 years ago
docker pull idio/jsonwikipedia