To be written.
Before running anything, download the dataset and put it into the dataset folder (name as "revPBE0-D3.data").
url="https://archive.materialscloud.org/record/file?filename=training-set.zip&file_id=4887cbec-a3b0-48f0-9e4b-1836c84a85b8&record_id=71";
curl -sSL $url -o training-set.zip && unzip training-set.zip &&\
mkdir -p datasets && mv training-set/input.data datasets/revPBE0-D3.data &&\
rm -r training-set*
The project contains Nextflow scripts and a Docker image which should make it easy to reproduce, the scripts used in are plain bash/python scripts.
For now the workflow is separated into three stages:
label.nf: computes DFTB+ labels for the entire datasettrain.nf: train the Delta-ML modelmdrun.nf: produces and analyzes the MD trajectoryTo run the project locally:
nextflow run label[train,mdrun].nf -resume
To run the project on a HPC cluster (assuming the queuing and project info in
nextflow.config are correct)
nextflow run label[train,mdrun].nf -profile rackham -resume
A Docker image containing all the requirements for this project is continuously built for this project, and the Nextflow script will use a singluarity image built out of it by default.
nextflow.config: config for running locally or on a HPC clusterpython/: Delta-ML calculator for ASEinputs
inputs: input files including that for PiNN and DFTBoutputs
datasets: downloaded/generated datasetsmodels: trained modelstrajs: MD trajectorieslabel.nf and generate the Delta-ML dataset.calculator for ASE/DeltaML, and run the MD.Content type
Image
Digest
Size
695.2 MB
Last updated
about 5 years ago
docker pull yqshao/deltaml