Experiments (VLDB 2021 candidate paper)
219
This repository contains all material used for the experiments of the following paper and is designed to enable anyone to reproduce them:
S. Irimescu, C. Berker Cikis, I. Müller, G. Fourny, G. Alonso. "Rumble: Data Independence for Large Messy Data Sets." In: PVLDB 14(4), 2020. DOI: 10.14778/3436905.3436910.
Each system used in the comparison has its own directory, all with the same structure: a subfolder with singlecore experiments and one with cluster experiments.
rumble/
cluster/
queries/
...
deploy.sh
upload.sh
run.sh
terminate.sh
singlecore/
queries/
...
deploy.sh
...
run_experiments.sh
zorba/
...
The flow for running the experiments is roughly the following:
deploy.sh.run_experiments.sh to run the desired subset of the configurations. Run run_experiments.sh.echo "$filelist" | upload.sh to upload the files stored in $filelist.cat queries/some-query.jq | run.sh` to run an individual query.terminate.sh.make -f path/to/common/make.mk -C results/results_date-of-experiment/ to parse the log files and produce result.jsonl with the statistics of all runs.We use a "prefix" of a sample for the single-core experiments and a "prefix" of the full data set for the cluster experiments. To download the sample or the full data set, use datasets/github/download-{sample,full}.sh.
We use the scripts of the original authors to download the data set and convert it to XML. Then we use datasets/vxquery-weather/convert.sh (which is based on a query by the original authors) to convert it to JSON.
The extract_prefix.sh and extract_prefix_s3.sh scripts help in producing a sub set of each data set and uploading it to S3.
Content type
Image
Digest
Size
151.4 MB
Last updated
almost 6 years ago
docker pull rumbledb/experiments-vldb21:vxquery-cli