+-----------------+ | SRA Database | | (SRRxxxxxxx) | +--------+--------+ | | fasterq-dump / SRA Toolkit v +-----------------+ | Kafka | | (fastq topics) | +--------+--------+ | | KafkaConsumer reads FASTQ streams v +-----------------+ | Alignment | | HISAT2 + BAM | +--------+--------+ | | FeatureCounts v +-----------------+ | Counts | | (gene counts) | +--------+--------+ | | Statistical Analysis (t-test) v +-----------------+ | DEGs | | (log2FC + p) | +--------+--------+ | | KEGG Enrichment (gseapy) v +-----------------+ | KEGG | | Pathway Maps | +-----------------+ Explanation of flow: SRA → Kafka: fasterq-dump streams reads directly into Kafka topics (per SRR ID). Kafka → Alignment: Consumers read FASTQ data from Kafka and feed it into HISAT2, producing BAM files. Alignment → Counts: featureCounts converts BAM alignments into gene-level counts. Counts → DEGs: Python calculates differential expression using t-tests, producing log2 fold changes and p-values.
DEGs → KEGG: Significant DEGs are sent to KEGG enrichment analysis with gseapy.
Content type
Image
Digest
sha256:47ca25a48…
Size
680.1 MB
Last updated
6 months ago
docker pull kartzz/ngss