Assessing the limits of genomic data integration for predicting protein networks

Abstract

Genomic data integration—the process of statistically combining diverse sources of information from functional genomics experiments to make large-scale predictions—is becoming increasingly prevalent. One might expect that this process should become progressively more powerful with the integration of more evidence. Here, we explore the limits of genomic data integration, assessing the degree to which predictive power increases with the addition of more features. We focus on a predictive context that has been extensively investigated and benchmarked in the past—the prediction of protein–protein interactions in yeast. We start by using a simple Naive Bayes classifier for integrating diverse sources of genomic evidence, ranging from coexpression relationships to similar phylogenetic profiles. We expand the number of features considered for prediction to 16, significantly more than previous studies. Overall, we observe a small, but measurable improvement in prediction performance over previous benchmarks, based on four strong features. This allows us to identify new yeast interactions with high confidence. It also allows us to quantitatively assess the inter-relations amongst different genomic features. It is known that subtle correlations and dependencies between features can confound the strength of interaction predictions. We investigate this issue in detail through calculating mutual information. To our surprise, we find no appreciable statistical dependence between the many possible pairs of features. We further explore feature dependencies by comparing the performance of our simple Naive Bayes classifier with a boosted version of the same classifier, which is fairly resistant to feature dependence. We find that boosting does not improve performance, indicating that, at least for prediction purposes, our genomic features are essentially independent. In summary, by integrating a few (i.e., four) good features, we approach the maximal predictive power of current genomic data integration; moreover, this limitation does not reflect (potentially removable) inter-relationships between the features.

Keywords

Classifier (UML)Naive Bayes classifierBoosting (machine learning)BiologyBayes' theoremMachine learningArtificial intelligenceGenomicsComputational biologyBayesian probabilityComputer scienceData miningGeneticsGenomeGeneSupport vector machine

MeSH Terms

AlgorithmsBayes TheoremComputational BiologyGenomicsProtein Interaction MappingROC Curve

Affiliated Institutions

Related Publications

Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy

Hanchuan Peng , Fuhui Long , Chen Ding

Feature selection is an important problem for pattern classification systems. We study how to select good features according to the maximal statistical dependency criterion base...

2005 IEEE Transactions on Pattern Analysis... 10050 citations

Additive logistic regression: a statistical view of boosting (With discussion and a rejoinder by the authors)

Jerome H. Friedman , Trevor Hastie , Robert Tibshirani

Boosting is one of the most important recent developments in\nclassification methodology. Boosting works by sequentially applying a\nclassification algorithm to reweighted versi...

2000 The Annals of Statistics 6819 citations

Dual Attention Network for Scene Segmentation

Jun Fu , Jing Liu , Haijie Tian +4 more

In this paper, we address the scene segmentation task by capturing rich contextual dependencies based on the self-attention mechanism. Unlike previous works that capture context...

2019 6497 citations

Molecular Similarity Searching Using Atom Environments, Information-Based Feature Selection, and a Naïve Bayesian Classifier

Andreas Bender , Hamse Y. Mussa , Robert C. Glen +1 more

A novel technique for similarity searching is introduced. Molecules are represented by atom environments, which are fed into an information-gain-based feature selection. A naïve...

2003 Journal of Chemical Information and C... 246 citations

Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project

Ewan Birney , J Stamatoyannopoulos , Anindya Dutta +97 more

We report the generation and analysis of functional data from multiple, diverse experiments performed on a targeted 1% of the human genome as part of the pilot phase of the ENCO...

2007 Nature 5179 citations

Publication Info

Year: 2005
Type: article
Volume: 15
Issue: 7
Pages: 945-953
Citations: 210
Access: Closed

External Links

Download PDF (Free) View on DOI.org PubMed Semantic Scholar

Social Impact

Altmetric

Assessing the limits of genomic data integration for predicting protein networks

PlumX Metrics

Social media, news, blog, policy document mentions

Citation Metrics

210

OpenAlex

Influential

159

CrossRef

Cite This

APA Style

                            
                                
                                    Long Lu, 
                                
                                    Yu Xia, 
                                
                                    Alberto Paccanaro
                                
                                et al.
                            
                            (2005). 
                            Assessing the limits of genomic data integration for predicting protein networks. 
                            Genome Research
                            , 15
                            (7)
                            , 945-953.
                            https://doi.org/10.1101/gr.3610305
                        

Identifiers

DOI: 10.1101/gr.3610305
PMID: 15998909
PMCID: PMC1172038

Data Quality

Data completeness: 86%