Sign inSign up

nfcore/bcellmagic

By nfcore

•Updated over 5 years ago

A preliminary container for the bcellmagic pipeline

Image
Machine learning & AI
Data science
0

4.7K

nfcore/bcellmagic repository overview

nf-core/bcellmagic

B cell repertoire analysis pipeline with the Immcantation framework..

GitHub Actions CI Status GitHub Actions Linting Status Nextflow DOI install with bioconda Docker Get help on Slack

⁠Introduction

The nf-core/bcellmagic pipeline is built to analyze B-cell repertoire sequencing data. It makes use of the Immcantation 2.5.0⁠ toolset and requires targeted sequencing data of the V, D, J and C regions of the B-cell receptor (primers for the V and C genes).

The pipeline is built using Nextflow⁠, a workflow tool to run tasks across multiple compute infrastructures in a very portable manner. It comes with docker containers making installation trivial and results highly reproducible.

⁠Quick Start

  1. Install nextflow⁠

  2. Install any of Docker⁠, Singularity⁠ or Podman⁠ for full pipeline reproducibility (please only use Conda⁠ as a last resort; see docs⁠)

  3. Download the pipeline and test it on a minimal dataset with a single command:

    nextflow run nf-core/bcellmagic -profile test,<docker/singularity/podman/conda/institute>
    

    Please check nf-core/configs⁠ to see if a custom config file to run nf-core pipelines already exists for your Institute. If so, you can simply use -profile <institute> in your command. This will enable either docker or singularity and set the appropriate execution settings for your local compute environment.

  4. Start running your own analysis!

    nextflow run nf-core/bcellmagic -profile <docker/singularity/podman/conda/institute> --input "metasheet_test.tsv" --cprimers "CPrimers.fasta" --vprimers "VPrimers.fasta"
    

See usage docs⁠ for all of the available options when running the pipeline.

⁠Pipeline Summary

By default, the pipeline currently performs the following steps:

  • Raw read quality control (FastQC)
  • Preprocessing (pRESTO)
    • Filtering sequences by sequencing quality.
    • Masking amplicon primers.
    • Pairing read mates.
    • Cluster sequences according to similarity, it helps identify if the UMI diversity was not high enough.
    • Building consensus of sequences with the same UMI barcode.
    • Re-pairing read mates.
    • Assembling R1 and R2 read mates.
    • Removing and annotating read duplicates with different UMI barcodes.
    • Filtering out sequences that do not have at least 2 duplicates.
  • Assigning gene segment alleles from teh IgBlast database (Change-O).
  • Determining the BCR / TCR genotype of the sample and finding the threshold for clone definition (TIgGER, SHazaM).
  • Clonal assignment: defining clonal lineages of the B-cell / T-cell populations (Change-O).
  • Reconstructing gene calls of germline sequences (Change-O).
  • Generating clonal trees (Alakazam).
  • Clonal analysis (Alakazam).
  • Repertoire comparison: calculation of clonal diversity and abundance (Alakazam).
  • Aggregating QC reports (MultiQC).

⁠Documentation

The nf-core/bcellmagic pipeline comes with documentation about the pipeline: usage⁠ and output⁠.

⁠Credits

nf-core/bcellmagic was originally written by Gisela Gabernet, Simon Heumos, Alexander Peltzer.

⁠Contributions and Support

If you would like to contribute to this pipeline, please see the contributing guidelines⁠.

For further information or help, don't hesitate to get in touch on the Slack #bcellmagic channel⁠ (you can join with this invite⁠).

⁠Citations

If you use nf-core/bcellmagic for your analysis, please cite it using the following doi: 10.5281/zenodo.3607408⁠

You can cite the nf-core publication as follows:

The nf-core framework for community-curated bioinformatics pipelines.

Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.

Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x⁠. ReadCube: Full Access Link⁠

In addition, references of tools and data used in this pipeline are as follows:

  • pRESTO Vander Heiden, J. A., Yaari, G., Uduman, M., Stern, J. N. H., O’Connor, K. C., Hafler, D. A., … Kleinstein, S. H. (2014). pRESTO: a toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. Bioinformatics, 30(13), 1930–1932. https://doi.org/10.1093/bioinformatics/btu138⁠.
  • SHazaM, Change-O Gupta, N. T., Vander Heiden, J. A., Uduman, M., Gadala-Maria, D., Yaari, G., & Kleinstein, S. H. (2015). Change-O: a toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data: Table 1. Bioinformatics, 31(20), 3356–3358. https://doi.org/10.1093/bioinformatics/btv359⁠.
  • Alakazam Stern, J. N. H., Yaari, G., Vander Heiden, J. A., Church, G., Donahue, W. F., Hintzen, R. Q., … O’Connor, K. C. (2014). B cells populating the multiple sclerosis brain mature in the draining cervical lymph nodes. Science Translational Medicine, 6(248). https://doi.org/10.1126/scitranslmed.3008879⁠.
  • TIgGER Gadala-maria, D., Yaari, G., Uduman, M., & Kleinstein, S. H. (2015). Automated analysis of high-throughput B-cell sequencing data reveals a high frequency of novel immunoglobulin V gene segment alleles. Proceedings of the National Academy of Sciences, 112(8), 1–9. https://doi.org/10.1073/pnas.1417683112⁠.
  • FastQC Download: https://www.bioinformatics.babraham.ac.uk/projects/fastqc/⁠
  • MultiQC Ewels, P., Magnusson, M., Lundin, S., & Käller, M. (2016). MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics , 32(19), 3047–3048. https://doi.org/10.1093/bioinformatics/btw354⁠. Download: https://multiqc.info/⁠.

Tag summary

Content type

Image

Digest

Size

1.4 GB

Last updated

over 6 years ago

docker pull nfcore/bcellmagic