NIH Comparative Genomics Resource Helps Improve Rigor of Genomic Studies
The NIH Comparative Genomics Resource (CGR) provides a centralized repository of tools and resources that aims to maximize the potential of eukaryotic research organisms and their genomic data for biomedical discovery. By bringing high-quality data and common tools together, CGR addresses some key challenges in genomics science, including inconsistencies in data quality and data siloed in separate systems. Moreover, this resource is another example of the wide array of NIH supported or managed data and biospecimen repositories that advance scientific rigor (see this Nexus article for more).
This resource from the NIH National Library of Medicine (NLM) reinforces NIH’s goals to strengthen our funded research, as outlined in NIH’s Gold Standard Science Plan. Its consistent data formats, access methods, and suite of interoperable tools make genomic data cleaner, easier to access and compare, and improve reusability. This toolkit enables scientists from many different fields within NIH’s mission to replicate, validate, and reanalyze earlier research, as well as generate their own hypotheses, develop robust models, and make discoveries that will improve human health. These capabilities make comparative analyses more accessible to all researchers—from computational researchers to scientists working at the bench or in clinical settings.
CGR allows researchers to:
- Find and download gene, transcript, protein, and genome sequences, annotation, and metadata through robust web and artificial intelligence-ready programmatic interfaces
- Infer functional and evolutionary relationships between sequences, as well as help identify members of gene families
- Graphically compare many assembled genomes at the whole genome, chromosome, or chromosome-region levels and identify genomic changes that could be significant to biology or evolution
- Generate a graphic display for nucleotide and protein sequence alignments to highlight sequence similarities and differences.
- Better understand sequence conservation with BLAST to the clusteredNR database, which provides fast results from a taxonomically diverse set of proteins
- Identify and remove contaminants in genome assemblies, reducing errors in analyses and conclusions and making the submission process faster
- Develop high-quality, submission-ready genome annotation for vertebrates, arthropods, and plants
A critical component of CGR is collaboration with the research community. CGR was designed to complement community-provided resources—including knowledge bases, genotype data, variation data, and image collections—by integrating their datasets with NLM’s centralized tools.
Selected Examples of How Researchers Are Using CGR
- Identified features that contribute to genomic instability across vertebrates using a survey of structural variants across 500 million years of evolution involving genome assemblies for more than 600 vertebrate species (see preprint)
- Studied the association of tumor prevalence with natural selection, improving our understanding of how genes affect cancer risk across species (see preprint)
- Provided core data access in the BRC Analytics platform for genomic analyses of pathogenic fungi and protists
Additional Resources:
Questions: [email protected]