Combining NLP and probabilistic categorisation for document and term selection for Swiss-Prot medical annotation.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 12855443.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Searching relevant publications for manual database annotation is a tedious task. In this paper, we apply a combination of Natural Language Processing (NLP) and probabilistic classification to re-rank documents returned by PubMed according to their relevance to Swiss-Prot annotation, and to identify significant terms in the documents. With a Probabilistic Latent Categoriser (PLC) we obtained 69% recall and 59% precision for relevant documents in a representative query. As the PLC technique provides the relative contribution of each term to the final document score, we used the Kullback-Leibler symmetric divergence to determine the most discriminating words for Swiss-Prot medical annotation. This information should allow curators to understand classification results better. It also has great value for fine-tuning the linguistic pre-processing of documents, which in turn can improve the overall classifier performance.
Medical subject headings
- Abstracting and Indexing
- Databases, Protein
- Models, Statistical
- Natural Language Processing
- Periodicals as Topic
- Proteins
- PubMed