Bayesian inference of protein-protein interactions from biological literature.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 19369495.
- Also identified by DOI 10.1093/bioinformatics/btp245 and PMC identifier 2732911.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Protein-protein interaction (PPI) extraction from published biological articles has attracted much attention because of the importance of protein interactions in biological processes. Despite significant progress, mining PPIs from literatures still rely heavily on time- and resource-consuming manual annotations. In this study, we developed a novel methodology based on Bayesian networks (BNs) for extracting PPI triplets (a PPI triplet consists of two protein names and the corresponding interaction word) from unstructured text. The method achieved an overall accuracy of 87% on a cross-validation test using manually annotated dataset. We also showed, through extracting PPI triplets from a large number of PubMed abstracts, that our method was able to complement human annotations to extract large number of new PPIs from literature. Programs/scripts we developed/used in the study are available at http://stat.fsu.edu/~jinfeng/datasets/Bio-SI-programs-Bayesian-chowdhary-zhang-liu.zip. Supplementary data are available at Bioinformatics online.
Medical subject headings
- Bayes Theorem
- Computational Biology
- Information Storage and Retrieval
- Protein Interaction Mapping
- Proteins