Sign inSign up

balercollaboration/baler

By balercollaboration

•Updated almost 3 years ago

This repository holds the docker image for the Baler data compression tool

Image
Machine learning & AI
Data science
2

311

balercollaboration/baler repository overview

⁠Introduction

This repository holds the docker image for the Baler data compression tool. Baler is developed as an open source project on GitHub⁠, published under the Apache-2.0 license⁠ togheter with the corresponding NOTICE file⁠

Baler is a tool used to test the feasibility of compressing different types of scientific data using machine learning-based autoencoders. Baler provides you with an easy way to:

  1. Train a machine learning model on your data
  2. Compress your data with that model. This will also save the compressed file and model
  3. Decompress the file using the model at a later time
  4. Plot the performance of the compression/decompression

⁠Getting Started

Follow the instructions on our GitHub repository to get started running Baler using Docker/Singularity/Apptainer⁠.

By using docker, executing the above 4 running modes uses one simple base command:

docker run \
-u ${UID}:${GID} \
--mount type=bind,source=${PWD}/projects/,target=/baler-root/projects \
--mount type=bind,source=${PWD}/data/,target=/baler-root/data \
pekman/baler:latest \
--project=example_CFD \
--mode=train

In order to train, compress, decompress, and plot change the "--mode" argument accordingly

⁠Contributing

If you wish to contribute, please see the contribution guidelines⁠.

⁠Project Details

For more information, visit our website⁠ (Work in Progress)

⁠Abstract

One common issue in vastly different fields of research and industry is the ever-increasing need for more data storage. With experiments taking more complex data at higher rates, the data recorded is quickly outgrowing the storage capabilities. This issue is very prominent in LHC experiments such as ATLAS where in five years the resources needed are expected to be many times larger than the storage available (assuming a flat budget model and current technology trends) [1]. Since the data formats used are already highly compressed, storage constraints could require more drastic measures such as lossy compression, where some data accuracy is lost during the compression process.

In our work, following from a number of undergraduate projects [2,3,4,5,6,7], we have developed an interdisciplinary open-source tool for machine learning-based lossy compression. The tool utilizes an autoencoder neural network, which is trained to compress and decompress data based on correlations between the different variables in the dataset. The process is lossy, meaning that the original data values and distributions cannot be reconstructed precisely. However, for certain variables and observables where the precision loss is tolerable, the high compression ratio allows for more data to be stored yielding greater statistical power.

The tool we have developed is called Baler and is available as an open source project [8][9].

[1] - https://cerncourier.com/a/time-to-adapt-for-big-data/⁠ [2] - http://lup.lub.lu.se/student-papers/record/9049610⁠ [3] - http://lup.lub.lu.se/student-papers/record/9012882⁠ [4] - http://lup.lub.lu.se/student-papers/record/9004751⁠ [5] - http://lup.lub.lu.se/student-papers/record/9075881⁠ [6] - https://zenodo.org/record/5482611#.Y3Yysy2l3Jz⁠ [7] - https://zenodo.org/record/4012511#.Y3Yyny2l3Jz⁠ [8] - https://zenodo.org/record/7817467#.ZED-65FBzmE⁠ [9] - https://github.com/baler-collaboration/baler⁠

Tag summary

Content type

Image

Digest

sha256:448dace29…

Size

658.1 MB

Last updated

almost 3 years ago

docker pull balercollaboration/baler:temp