A simple and practical dictionary-based approach for identification of proteins in Medline abstracts.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 14764613.
- Also identified by PMC identifier 400515.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The aim of this study was to develop a practical and efficient protein identification system for biomedical corpora. The developed system, called ProtScan, utilizes a carefully constructed dictionary of mammalian proteins in conjunction with a specialized tokenization algorithm to identify and tag protein name occurrences in biomedical texts and also takes advantage of Medline "Name-of-Substance" (NOS) annotation. The dictionaries for ProtScan were constructed in a semi-automatic way from various public-domain sequence databases followed by an intensive expert curation step. The recall and precision of the system have been determined using 1000 randomly selected and hand-tagged Medline abstracts. The developed system is capable of identifying protein occurrences in Medline abstracts with a 98% precision and 88% recall. It was also found to be capable of processing approximately 300 abstracts per second. Without utilization of NOS annotation, precision and recall were found to be 98.5% and 84%, respectively. The developed system appears to be well suited for protein-based Medline indexing and can help to improve biomedical information retrieval. Further approaches to ProtScan's recall improvement also are discussed.
Medical subject headings
- Information Storage and Retrieval
- MEDLINE
- Proteins
- Terminology as Topic