Sign inSign up

gfinak/abstar

By gfinak

•Updated almost 6 years ago

The abstar tool from briney/abstar, built with python 3

Image
0

10K+

gfinak/abstar repository overview

⁠VISC_eODGT8_abstar_pipeline

A dockerized nextflow pipeline for processing B cell sequencing using abstar for the eODGT8 study.

⁠Cloning the repo from github

The repository holds all the code to build the pipeline from scratch. The cloned repo can be used to build the docker container with the pipeline code, or to run the pipeline locally (without using docker). The latter is good for testing and development.

To clone the repo:

git clone github.com:FredHutch/VISC_eODGT8_abstar_pipeline .

The repository will be placed in a local directory named VISC_eODGT8_abstar_pipeline by default.

⁠Installing docker

In order to use build or use a docker container of the pipeline you must have docker⁠ installed and running.

⁠Building the docker container

You can build the docker container from the repo by entering:

make nextflow

⁠Using a pre-built container

Rather than building the container yourself, you can use the prebuilt one hosted on dockerhub.

⁠Pulling the docker container from dockerhub⁠.

If you wish to run the pipeline using the pre-built docker container you will need to pull it from dockerhub.

docker pull gfinak/abstar:nextflow

This will pull the pre-build contianer.

You can then run the pipeline using nextflow⁠

⁠Installing nextflow⁠

  1. To install nextflow, you need java 8 or later.

java -version

  1. Install nextflow (this will install in the current directory, /usr/local/bin is recommended).

curl -s https://get.nextflow.io | bash

  1. Set your PATH to point to the location of the nextflow binary.

  2. Run "Hello world" to confirm it is working

./nextflow run hello

⁠Running the G001 sequencing pipeline using nextflow

⁠The simplest way, using the pre-built docker container.

From the root of the cloned repository, run:

`nextflow run main.nf --indir /path/to/ab1/files/

Note the trailing slash on the path.

The pipeline will run on all ab1 files in the provided path (with recursion depending on the glob), using the pre-built dockerized version of the pipeline tools.

Alternately, if you want to use the local files, rather than the docker container, then you can add -without-docker to the command line:

nextflow run main.nf --indir /path/to/ab1/files/ -without-docker

Again note the trailing slash on the path.

⁠Nextflow Pipeline Configuration

The configuration for this nextflow pipeline is in nextflow.config. There you can see it is using a docker image named gfinak/abstar:nextflow by default. This is what is pulled from dockerhub.

⁠Making changes

If you make changes to the pipeline, you will have to run the pipeline using the -without-docker option. This will run it using your local modified files.

If you're done testing, make a pull request and push your changes to github. When they're accepted, the docker container will re-built and pushed to dockerhub.

Alternately buld your own container, push it to docker hub and update the nextflow.config to reflect the different container instance.

⁠The more complicated way: without cloning the repo.

Nextflow is built around pipeline sharing. To that end, it works best with public repositories.

You can have nextflow pull everything from the remote github repo.

To run the pipeline, you do the following:

nextflow run FredHutch/VISC_eODGT8_abstar_pipeline -with-docker gfinak/abstar:nextflow --indir /path/to/ab1/files/

Here we're specifying the github repository hosting the pipeline directly rather than the pipeline .nf nextflow script. We are not even cloning the repository. Nextflow will take care of this, pulling the repository, pulling the container and running.

Nextflow will search for the docker container named gfinak:/abstar:nextflow on dockerhub and will search for a main.nf file and a nextflow.config file at the github repository under FredHutch/VISC_eODGT8_abstar_pipeline. It clones the repo locally, storing it in $HOME/.nextflow/assets and runs the pipeline using the docker container.

⁠Where are the results?

In both these cases, output will be generated in a directory named results at the current working directory where you launched the pipeline.

Temporary files will be placed in a directory called work.

⁠Building the container from scratch.

You may, but don’t have to, build the container from scratch.

If you've cloned the repository locally, the Makefile contains processes for building the G001 pipeline docker container via

make nextflow

⁠Resuming

You can resume a partial run of the pipeline by adding -resume to the command line:

For example:

nextflow run main.nf --indir /path/to/ab1/files -resume

This will use any cached results in the work directory work.

⁠Next steps

QC reports based on the September 19th 2019 meeting in DC will be added to the pipeline.

Sequencing data sets will be output with all the required contents, again based on the output of that meeting.

TODO: Process sequencing manifests TODO: Add control annotations. TODO: filter empty wells (wells where there are no cells, but not NTC wells). TODO: Identify NTC wells with sequence and filter other wells with those sequences. TODO: Identify replicate sequences at the nucleotide level. TODO: Output QC reports with per-plate summaries based on Lexi's and Lamar's recommendations.

⁠QC Reports

Currently one QC report is generated, results are found in results/QC_report.html. this report is currently a live document. It's contents change as we figure out what we need to put in there. This report shows

  • The total number of sequences.
  • The number of uploaded duplicate sequences.
  • The number of unique sequences.
  • The number of sequences that were dropped in favor of a higher-scoring replicate PCR.
  • The number of sequences that were renamed in some way due to invalid naming schemes.
  • The number that have an invalid visit format arising from multi-visit plates.
  • The number that were not found mapping to the lab's sequences manifest.
  • The plates that can't be matched to the manifest based on the tracking number.
  • The plates that have discordant metadata between the manifest and the received data based on the tracking number.
  • Well level QC showing the number of wells that don't match the manifest (e.g. should be empty according to the manifest, but for which we have data).
    • The well level QC is done for paired, non-paired sequences.
  • The per-plate QC showing the number of heavy, lambda, and kappa sequences that were found to be paired or unpaired, productive or non-productive.
  • The productive sequences are output to file.

⁠Other files

There are a number of other files in the output directory.

⁠Stage 1 files.
  • SEQUENCE_EXCEPTIONS_LOG.csv: A list of duplicate sequence files based on metadata and tracking number.
  • SEQUENCE_RENAMING_LOG.csv: A list of sequences and how they were renamed to match the expected naming format at the first stage of data cleaining.
  • Invalid_visit_in_sequence_id.csv: A list of sequences with invalid visit IDs after the first stage of data cleaning.
  • MANIFEST_TRACKING_NOT_IN_RECEIVE_SEQS.csv: A list of plates by tracking number that were not received, but should be according to the manifest.
  • RECEIVED_SEQUENCE_TRACKING_NOT_IN_MANIFEST.csv: A list of plates by tracking number that were received but are not listed as "sent to visc" in the manifest.
  • duplicate_sequence_files.csv: Duplicate sequence files based on metadata and tracking number, used for reporting.
⁠Stage 2 files.
  • FHCRC_Seq_Manifest_G001_Log_for_VISC_YYYYMMDD_vXX.csv: The expanded FHCRC sample manifest with information on wells with cells, controls, etc.
  • VRC_Sequencing_Manifest_YYYYMMDD_vXX.csv: The expanded VRC sample manifest with information on wells with cells, controls, etc.
  • CONVERSION_ERRORS.csv: A list of transformations done to select optimal scoring sequences or to annotate sequences after abstar is run.
  • NOT_FOUND.csv: Sequences not found in the manifest at the well level after abstar is run.
  • non_productive_sequences_YYYY-MM-DD.json: Non-productive sequences based on abstar.
  • output_paired_YYYY-MM-DD.csv: Productive paired sequences after abstar, merged with manifest metadata.
  • output_unpaired_YYYY-MM-DD.csv: Unpaired productive sequences after abstar, merged with manifest metadata.
⁠QC and output files.
  • QC_sequences.html: The QC report.
  • productive_sequences.csv: All productive sequences, merged with manifest metdata, including VRC01 class calls based on VH gene usage and CDRL3 length.

Tag summary

Content type

Image

Digest

Size

917.9 MB

Last updated

almost 6 years ago

docker pull gfinak/abstar