MCAAT - Metagenomic CRISPR Array Analysis Tool
283
Finds CRISPR arrays in raw, un-assembled metagenomic reads. Builds a succinct de Bruijn graph and detects multicycles - the structural signature of CRISPR repeat-spacer arrays - without any prior assembly step.
The Docker image is based on debian:bookworm-slim and ships only the mcaat binary and its runtime dependencies.
docker pull feeka94/mcaat:1.0.1
Mount the directory containing your reads and specify paths inside the container under /data:
docker run --rm -v $(pwd):/data feeka94/mcaat:1.0.1 \
--input-files /data/reads_R1.fastq /data/reads_R2.fastq \
--output-folder /data/results
For paired-end reads:
docker run --rm -v $(pwd):/data feeka94/mcaat:1.0.1 \
--input-files /data/reads_R1.fastq /data/reads_R2.fastq \
--output-folder /data/results
For a single-end file:
docker run --rm -v $(pwd):/data feeka94/mcaat:1.0.1 \
--input-files /data/reads.fastq \
--output-folder /data/results
From a pre-built graph (skips graph construction):
docker run --rm -v $(pwd):/data feeka94/mcaat:1.0.1 \
--graph /data/graph \
--output-folder /data/results
Required (one of):
| Flag | Description |
|---|---|
--input-files <file1> [file2] | One or two FASTA/FASTQ files — plain or gzipped. One file = single-end, two = paired-end |
--graph <path> | Pre-built SDBG graph directory from a previous run (skips graph construction) |
Optional:
| Flag | Default | Description |
|---|---|---|
--output-folder <path> | mcaat_run_YYYY-MM-DD_HH-MM-SS/ | Output directory |
--ram <amount> | 95% of system RAM | Memory cap. Units: B, K, M, G (e.g. --ram 8G) |
--threads <num> | CPU cores − 2 | Thread count |
--cycle-max-length <int> | 77 | Maximum cycle length to search |
--cycle-min-length <int> | 27 | Minimum cycle length to search |
--threshold-multiplicity <int> | 20 | Min edge multiplicity for cycle start nodes |
--low-abundance <true|false> | true | Enable low-abundance mode |
--settings <path> | — | Key=value settings file (CLI flags override it) |
--benchmark <file> | — | File with expected CRISPR sequences (one per line) for evaluation |
--help, -h | — | Show usage and exit |
Pass a key=value file with --settings. CLI flags override any value from the file.
input-files=/data/R1.fastq /data/R2.fastq
ram=128G
threads=26
output-folder=/data/results
cycle-max-length=77
cycle-min-length=27
threshold-multiplicity=20
low-abundance=true
docker run --rm -v $(pwd):/data feeka94/mcaat:1.0.1 --settings /data/settings.txt
<output-folder>/
├── CRISPR_Arrays_1.txt # detected arrays (split into numbered files if large)
├── graph/ # succinct de Bruijn graph files
└── cycles/ # raw cycle data
Each CRISPR_Arrays_N.txt file has a short header followed by one block per array:
# MCAAT — CRISPR Array Output
# Generated : 2026-05-12 10:30:21
# Arrays : 42
# Spacers : 312
>Array_1 spacers=8
ATCGATCGATCGATCGATCGATCG
-------------------- AACCCGGTTAATCGATCGTTTCGAGC
-------------------- TTGGCCAATCGATCGATCAAAACGGG
ATCGATCGATCGATCTATCG GGAATTCCAATCGATCGAATACCCAC ← repeat variant
The consensus repeat sequence is on its own line. Each spacer entry shows the repeat variant (or dashes when it matches the consensus exactly) followed by the spacer sequence.
# Remove the image
docker rmi feeka94/mcaat:1.0.1
# If the image has dependent containers, remove those first
docker rm $(docker ps -aq --filter ancestor=feeka94/mcaat:1.0.1)
docker rmi feeka94/mcaat:1.0.1
# Or force-remove everything at once
docker rmi -f feeka94/mcaat:1.0.1
To also free up dangling/unused layers afterwards:
docker image prune
If you use MCAAT please cite: https://academic.oup.com/microlife/article/doi/10.1093/femsml/uqaf016/8205558
Please write an issue on our GitHub page if any problems occur.
Content type
Image
Digest
sha256:9426a4889…
Size
30.5 MB
Last updated
2 months ago
docker pull feeka94/mcaat:1.0.1