NCBI Reference Sequences: current status, policy and new initiatives

Kim D. Pruitt; Tatiana Tatusova; William Klimke; Donna Maglott

doi:10.1093/nar/gkn721

Abstract

NCBI's Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. RefSeq records integrate information from multiple sources and represent a current description of the sequence, the gene and sequence features. The database includes over 5300 organisms spanning prokaryotes, eukaryotes and viruses, with records for more than 5.5 x 10(6) proteins (RefSeq release 30). Feature annotation is applied by a combination of curation, collaboration, propagation from other sources and computation. We report here on the recent growth of the database, recent changes to feature annotations and record types for eukaryotic (primarily vertebrate) species and policies regarding species inclusion and genome annotation. In addition, we introduce RefSeqGene, a new initiative to support reporting variation data on a stable genomic coordinate system.

Keywords

RefSeqAnnotationBiologyGenomeEnsemblSequence (biology)Gene AnnotationComputational biologyGenome projectReference genomeFeature (linguistics)BioinformaticsGenomicsInformation retrievalGeneDatabaseGeneticsComputer science

Affiliated Institutions

Related Publications

RefSeq: an update on mammalian reference sequences

Kim D. Pruitt , Garth Brown , Susan M. Hiatt +26 more

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database is a collection of annotated genomic, transcript and protein sequence records deriv...

2013 Nucleic Acids Research 994 citations

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Kim D. Pruitt , Tatiana Tatusova , D. R. Maglott

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and protei...

2006 Nucleic Acids Research 4555 citations

NCBI Reference Sequences (RefSeq): current status, new features and genome annotation policy

Kim D. Pruitt , Tatiana Tatusova , Garth Brown +1 more

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database is a collection of genomic, transcript and protein sequence records. These records ...

2011 Nucleic Acids Research 1166 citations

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Kim D. Pruitt

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) provides a non-redundant collection of sequen...

2004 Nucleic Acids Research 1622 citations

Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation

Nuala A. O’Leary , Matt W. Wright , J. Rodney Brister +52 more

The RefSeq project at the National Center for Biotechnology Information (NCBI) maintains and curates a publicly available database of annotated genomic, transcript, and protein ...

2015 Nucleic Acids Research 6668 citations

Publication Info

Year: 2008
Type: article
Volume: 37
Issue: Database
Pages: D32-D36
Citations: 740
Access: Closed

External Links

View on DOI.org

Social Impact

Altmetric

NCBI Reference Sequences: current status, policy and new initiatives

PlumX Metrics

Social media, news, blog, policy document mentions

Citation Metrics

740

OpenAlex

Cite This

APA Style

                            
                                
                                    Kim D. Pruitt, 
                                
                                    Tatiana Tatusova, 
                                
                                    William Klimke
                                
                                et al.
                            
                            (2008). 
                            NCBI Reference Sequences: current status, policy and new initiatives. 
                            Nucleic Acids Research
                            , 37
                            (Database)
                            , D32-D36.
                            https://doi.org/10.1093/nar/gkn721
                        

Identifiers

DOI: 10.1093/nar/gkn721