Pull the Docker image available from DockerHub:
docker pull mazzalab/fastqwiper
Once downloaded the image, you can type:
docker run --rm -ti --name fastqwiper -v "YOUR_LOCAL_PATH_TO_DATA_FOLDER:/fastqwiper/data" mazzalab/fastqwiper paired 8 sample 50000000 33 ACGTN
where:
- YOUR_LOCAL_PATH_TO_DATA_FOLDER is the path to the folder where the fastq.gz files to be wiped are located;
- paired triggers the cleaning of R1 and R2. Alternatively, single will trigger the wiping of individual FASTQ files;
- 8 is the number of computing cores to be spawned;
- sample is part of the names of the FASTQ files to be wiped. Be aware that: for paired-end files (e.g., "sample_R1.fastq.gz" and "sample_R2.fastq.gz"), your files must finish with _R1.fastq.gz and _R2.fastq.gz. Therefore, the argument to pass is everything before these texts: sample in this case. For single end/individual files (e.g., "excerpt_R1_001.fastq.gz"), your file must end with the string .fastq.gz; the preceding text, i.e., excerpt_R1_001 in this case, will be the text to be passed to the command as an argument. E.g.,
- 50000000 (optional) is the number of rows-per-chunk (used when cores>1. It must be a number multiple of 4). Increasing this number too much would reduce the parallelism advantage. Decreasing this number too much would increase the number of chunks more than the number of available cpus, making parallelism inefficient. Choose this number wisely depending on the total number of reads in your starting file.
- 33 (optional) is the ASCII offset (33=Sanger, 64=old Solexa)
- ACGTN (optional) is the allowed alphabet in the SEQ line of the FASTQ file