Sign inSign up

getwilds/clair3

Sponsored OSS

By Fred Hutch Data Science Lab

•Updated 6 months ago

Docker image for Clair3 in Fred Hutch OCDO's WILDS

Image
0

915

getwilds/clair3 repository overview

⁠Clair3

This directory contains Docker images for Clair3, a germline small variant caller optimized for long-read sequencing data (Oxford Nanopore and PacBio HiFi) that identifies SNPs and indels using a pileup-plus-full-alignment approach with PyTorch-based deep learning models.

⁠Available Versions

⁠Image Details

These Docker images are built from nvidia/cuda:12.6.3-runtime-ubuntu24.04 with Miniforge installed on top, and include:

  • Clair3 v2.0.0: Germline small variant caller for long-read sequencing (ONT, PacBio HiFi), built from source
  • samtools: BAM/CRAM file processing
  • whatshap: Read-based phasing support
  • PyTorch with CUDA 12.6 support: GPU-accelerated deep learning inference (pass --use_gpu to enable)
  • PyPy 3.11 v7.3.20: Preprocessing acceleration
  • longphase v1.7.3: Long-read phasing

The images support both CPU and GPU execution — GPU mode is opt-in via the --use_gpu flag, so the image works in CPU-only environments without any changes. All 14 pre-trained sequencing models are bundled in the image at /opt/models/.

⁠Citation

If you use Clair3 in your research, please cite the original authors:

Zheng, Z., Li, S., Su, J., Leung, A.W.S., Lam, T.W., & Luo, R. (2022).
Symphonizing pileup and full-alignment for deep learning-based long-read variant calling.
Nature Computational Science, 2, 797–803.
https://doi.org/10.1038/s43588-022-00387-x

Tool homepage: https://github.com/HKU-BAL/Clair3⁠

⁠Usage

⁠Docker
# Pull the latest version
docker pull getwilds/clair3:latest

# Or pull a specific version
docker pull getwilds/clair3:2.0.0

# Alternatively, pull from GitHub Container Registry
docker pull ghcr.io/getwilds/clair3:latest
⁠Singularity/Apptainer
# Pull the latest version
apptainer pull docker://getwilds/clair3:latest

# Or pull a specific version
apptainer pull docker://getwilds/clair3:2.0.0

# Alternatively, pull from GitHub Container Registry
apptainer pull docker://ghcr.io/getwilds/clair3:latest
⁠Example Commands

All 14 pre-trained models are bundled in the image at /opt/models/. Pass the appropriate model path for your sequencing platform via --model_path.

# Run Clair3 on ONT R10.4.1 data (CPU mode)
docker run --rm \
  -v /path/to/data:/data \
  getwilds/clair3:latest \
  run_clair3.sh \
  --bam_fn=/data/sample.bam \
  --ref_fn=/data/reference.fa \
  --threads=4 \
  --platform=ont \
  --model_path=/opt/models/r1041_e82_400bps_sup_v500 \
  --output=/data/clair3_output

# Run Clair3 on ONT R10.4.1 data with GPU acceleration (requires NVIDIA Container Toolkit)
docker run --rm --gpus all \
  -v /path/to/data:/data \
  getwilds/clair3:latest \
  run_clair3.sh \
  --bam_fn=/data/sample.bam \
  --ref_fn=/data/reference.fa \
  --threads=4 \
  --platform=ont \
  --model_path=/opt/models/r1041_e82_400bps_sup_v500 \
  --output=/data/clair3_output \
  --use_gpu

# Run Clair3 on PacBio HiFi data (CPU mode)
docker run --rm \
  -v /path/to/data:/data \
  getwilds/clair3:latest \
  run_clair3.sh \
  --bam_fn=/data/sample.hifi.bam \
  --ref_fn=/data/reference.fa \
  --threads=8 \
  --platform=hifi \
  --model_path=/opt/models/hifi_revio \
  --output=/data/clair3_hifi_output

# Run using Apptainer with GPU
apptainer run --nv \
  --bind /path/to/data:/data \
  docker://getwilds/clair3:latest \
  run_clair3.sh \
  --bam_fn=/data/sample.bam \
  --ref_fn=/data/reference.fa \
  --threads=4 \
  --platform=ont \
  --model_path=/opt/models/r1041_e82_400bps_sup_v500 \
  --output=/data/clair3_output \
  --use_gpu

⁠Important Notes

⁠GPU Execution

GPU mode is opt-in via --use_gpu and requires the NVIDIA Container Toolkit⁠ on the host. Pass --gpus all to Docker or --nv to Apptainer. Without these flags the image runs in CPU mode regardless of available hardware. The image uses CUDA 12.6, which is compatible with NVIDIA driver ≥525.60.13.

⁠Platform Support

These images are built for linux/amd64 only.

⁠Dockerfile Structure

The Dockerfile follows these main steps:

  1. Uses nvidia/cuda:12.6.3-runtime-ubuntu24.04 as the base image
  2. Adds metadata labels for documentation and attribution
  3. Installs Miniforge and creates a conda environment with build tools and bioinformatics dependencies (samtools, whatshap, parallel, etc.)
  4. Installs PyTorch with CUDA 12.6 support and Python dependencies (numpy, h5py, torchmetrics, etc.) via uv
  5. Installs Clair3 from source
  6. Installs PyPy 3.11 v7.3.20 for preprocessing acceleration
  7. Downloads all 14 pre-trained PyTorch models to /opt/models/
  8. Runs run_clair3.sh --version and a PyTorch import check as smoke tests

⁠Security Scanning and CVEs

These images are regularly scanned for vulnerabilities using Docker Scout. However, due to the nature of bioinformatics software and their dependencies, some Docker images may contain components with known vulnerabilities (CVEs).

Use at your own risk: While we strive to minimize security issues, these images are primarily designed for research and analytical workflows in controlled environments.

For the latest security information about this image, please check the CVEs_*.md files in this directory⁠, which are automatically updated through our GitHub Actions workflow. If a particular vulnerability is of concern, please file an issue⁠ in the GitHub repo citing which CVE you would like to be addressed.

⁠Source Repository

These Dockerfiles are maintained in the WILDS Docker Library⁠ repository.

Tag summary

Content type

Image

Digest

sha256:246a5a91f…

Size

5.8 GB

Last updated

6 months ago

docker pull getwilds/clair3