The abstar tool from briney/abstar, built with python 3
10K+
A dockerized nextflow pipeline for processing B cell sequencing using abstar for the eODGT8 study.
The repository holds all the code to build the pipeline from scratch. The cloned repo can be used to build the docker container with the pipeline code, or to run the pipeline locally (without using docker). The latter is good for testing and development.
To clone the repo:
git clone github.com:FredHutch/VISC_eODGT8_abstar_pipeline .
The repository will be placed in a local directory named VISC_eODGT8_abstar_pipeline by default.
In order to use build or use a docker container of the pipeline you must have docker installed and running.
You can build the docker container from the repo by entering:
make nextflow
Rather than building the container yourself, you can use the prebuilt one hosted on dockerhub.
If you wish to run the pipeline using the pre-built docker container you will need to pull it from dockerhub.
docker pull gfinak/abstar:nextflow
This will pull the pre-build contianer.
You can then run the pipeline using nextflow
java -version
/usr/local/bin is recommended).curl -s https://get.nextflow.io | bash
Set your PATH to point to the location of the nextflow binary.
Run "Hello world" to confirm it is working
./nextflow run hello
From the root of the cloned repository, run:
`nextflow run main.nf --indir /path/to/ab1/files/
Note the trailing slash on the path.
The pipeline will run on all ab1 files in the provided path (with recursion depending on the glob), using the pre-built dockerized version of the pipeline tools.
Alternately, if you want to use the local files, rather than the docker container, then you can add -without-docker to the command line:
nextflow run main.nf --indir /path/to/ab1/files/ -without-docker
Again note the trailing slash on the path.
The configuration for this nextflow pipeline is in nextflow.config.
There you can see it is using a docker image named gfinak/abstar:nextflow by default.
This is what is pulled from dockerhub.
If you make changes to the pipeline, you will have to run the pipeline using the -without-docker option. This will
run it using your local modified files.
If you're done testing, make a pull request and push your changes to github. When they're accepted, the docker container will re-built and pushed to dockerhub.
Alternately buld your own container, push it to docker hub and update the nextflow.config to reflect the different container instance.
Nextflow is built around pipeline sharing. To that end, it works best with public repositories.
You can have nextflow pull everything from the remote github repo.
To run the pipeline, you do the following:
nextflow run FredHutch/VISC_eODGT8_abstar_pipeline -with-docker gfinak/abstar:nextflow --indir /path/to/ab1/files/
Here we're specifying the github repository hosting the pipeline directly rather than the pipeline .nf nextflow script. We are not even cloning the repository. Nextflow will take care of this, pulling the repository, pulling the container and running.
Nextflow will search for the docker container named gfinak:/abstar:nextflow on dockerhub and will search for a main.nf file and a nextflow.config file at the github repository under FredHutch/VISC_eODGT8_abstar_pipeline. It clones the repo locally, storing it in $HOME/.nextflow/assets and runs the pipeline using the docker container.
In both these cases, output will be generated in a directory named results at the current working directory where you launched the pipeline.
Temporary files will be placed in a directory called work.
You may, but don’t have to, build the container from scratch.
If you've cloned the repository locally, the Makefile contains processes for building the G001 pipeline docker container via
make nextflow
You can resume a partial run of the pipeline by adding -resume to the command line:
For example:
nextflow run main.nf --indir /path/to/ab1/files -resume
This will use any cached results in the work directory work.
QC reports based on the September 19th 2019 meeting in DC will be added to the pipeline.
Sequencing data sets will be output with all the required contents, again based on the output of that meeting.
TODO: Process sequencing manifests TODO: Add control annotations. TODO: filter empty wells (wells where there are no cells, but not NTC wells). TODO: Identify NTC wells with sequence and filter other wells with those sequences. TODO: Identify replicate sequences at the nucleotide level. TODO: Output QC reports with per-plate summaries based on Lexi's and Lamar's recommendations.
Currently one QC report is generated, results are found in results/QC_report.html.
this report is currently a live document. It's contents change as we figure out what we need to put in there.
This report shows
There are a number of other files in the output directory.
Content type
Image
Digest
Size
917.9 MB
Last updated
almost 6 years ago
docker pull gfinak/abstar