This container will tar and plot your data on the PromethION device before transferring.
905
For data-handling on the PromethION beta device.






There is no need to repeat what the internet has already provided Click here to see a nice straight forward guide to downloading docker.
You may need to log out and log back in again for the super user permissions to take place. If that is not feasible, you may need to precede the docker commands below with sudo
docker pull alexiswl/betaduck:18.07.1-3
18.07.1-3 represents the current MinKNOW version.
First we'll generate the config file. This is used as an input in the next command as to which files and folders to work on.
For assistance with options use
docker run alexiswl/betaduck config --help
betaduck config parameters
/data/basecalled/<sample>/<flowcell_port>//data/basecalled/<sample>/<flowcell_port>//data/basecalled/<sample>/<flowcell_port>/reads/data/basecalled/<sample>/<flowcell_port>/config.yamldocker parameters
You will also need to bind the /data volume to the container when executing the script.
Example:
docker run --volume /data:/data alexiswl/betaduck config ..parameters
Most of the hard work has been done for us in the previous script.
betaduck tidy parameters
docker parameters
Here is where docker shines, it can restrict the cpus and memory utilisations of a given container as to not blow up your system.
Having said that, the PromethION beta device is pretty powerful machine. It has 96 cpus with 400 Gb of memory and all of this on SSD drives.
Use the top and free commands to view the current utilisations of your system.
The --cpus parameter should be 1 higher than the --threads parameter used by betaduck.
Example:
docker run --memory=50g --cpus=7 --volume=/data:/data alexiswl/betaduck tidy ..parameters
Now we get our rewards, some plots produced from the seaborn and matplotlib libraries
betaduck plot parameters
/data/basecalled/<sample>/<flowcell_port>/sequencing_summary//data/basecalled/<sample>/<flowcell_port>/fastq/data/basecalled/<sample>/<flowcell_port>/plotsNow that your nanopore data is tidy, you can rsync the data across using rsync. You will need to specify which files to include and exclude in the rsync comamnd.
rsync --archive ' \ # Archive allows rsync to find files recursively
--include='*/' \ # Look for the following file endinds in subfolders
--include='*.fast5.tar.gz' \ # Find all the fast5 tars
--include='*.fastq.gz' \ # Find all the gzipped fastq files
--include='*.sequencing_summary.txt' \ # Find all the moved sequencing summary files
--include='*.png' \ # Carry over any plots that have been generated
--exclude='*' \ # Exclude all other files
--prune-empty-dirs \ # Don't download folders that won't have files in them (like reads/0 etc)
/path/to/data/basecalled/reads
/dest/directory
This is also acceptable to do whilst the run is still being generated or the betaduck tidy script is running. fast5.tar.gz files are written as fast5.tar.gz.tmp initially and then moved so they won't be listed by the rsync if they're still being written to.
--remove-source-files if you wish to remove the data from the PromethION device.
docker run --volume /data:/data --volume /etc/localtime:/etc/localtime alexiswl/betaduck ..parametersContent type
Image
Digest
Size
2 GB
Last updated
over 7 years ago
docker pull alexiswl/betaduck