OpenCGA provides a scalable and high-performance platform for big data analysis and visualization
10K+
OpenCGA (https://github.com/opencb/opencga) provides a scalable and high-performance solution for big data analysis and visualization in a shared environment. OpenCGA integrates some of the OpenCB projects and implements, in addition, other components: i) a storage engine framework to store and index alignments and genomic variants into different NoSQL such as MongoDB or Hadoop HBase - the current implementation can store efficiently thousands of gVCF files while remaining responsive when querying data; ii) a Catalog which keeps track of users, projects, files, samples, annotations, etc and also provides authentication and authorization capabilities; iii) analysis engine to execute genomic analysis in a traditional HPC cluster or in Hadoop. OpenCGA has implemented a command line and a RESTful web services to manage and query all the data.
Content type
Image
Digest
Size
712.9 MB
Last updated
over 6 years ago
docker pull opencb/opencga:f0b252daf