Rapid determination of appropriate reference genomes.
1.0K
ReferenceSeeker determines closely related reference genomes from RefSeq (https://www.ncbi.nlm.nih.gov/refseq) following a scalable hierarchical approach combining an ultra-fast kmer-based database lookup of candidate reference genomes and subsequent computation of specific average nucleotide identity (ANI) values for the rapid selection of suitable reference genomes.
ReferenceSeeker computes kmer-based genome distances between a query genome and and a database built on RefSeq genomes via Mash (Ondov et al. 2016). Therefore, only complete genomes or those stated as 'representative' or 'reference' genome are included. ReferenceSeeker offers pre-built databases for a broad spectrum of microbial taxonomic groups, i.e. bacteria, archaea, fungi, protozoa and viruses. For resulting candidates ReferenceSeeker subsequently computes ANI values picking genomes meeting community standard thresholds (ANI >= 95 % & conserved DNA >= 69 %) (Goris, Konstantinos et al. 2007) ranked by ANI and conserved DNA. Additionally, ReferenceSeeker can use MeDuSa (Bosi, Donati et al. 2015) to scaffold contigs based on the 20 closest reference genomes.
Source code, download links for related pre-built databases and further information are available on GitHub: https://github.com/oschwengers/referenceseeker
Content type
Image
Digest
Size
298.1 MB
Last updated
over 7 years ago
docker pull oschwengers/referenceseeker