A simple and practical dictionary-based approach for identification of proteins in Medline abstracts.

Egorov, Sergei; Yuryev, Anton; Daraselia, Nikolai · J Am Med Inform Assoc · 2004

other · Level V

Where this comes from

Abstract

The aim of this study was to develop a practical and efficient protein identification system for biomedical corpora. The developed system, called ProtScan, utilizes a carefully constructed dictionary of mammalian proteins in conjunction with a specialized tokenization algorithm to identify and tag protein name occurrences in biomedical texts and also takes advantage of Medline "Name-of-Substance" (NOS) annotation. The dictionaries for ProtScan were constructed in a semi-automatic way from various public-domain sequence databases followed by an intensive expert curation step. The recall and precision of the system have been determined using 1000 randomly selected and hand-tagged Medline abstracts. The developed system is capable of identifying protein occurrences in Medline abstracts with a 98% precision and 88% recall. It was also found to be capable of processing approximately 300 abstracts per second. Without utilization of NOS annotation, precision and recall were found to be 98.5% and 84%, respectively. The developed system appears to be well suited for protein-based Medline indexing and can help to improve biomedical information retrieval. Further approaches to ProtScan's recall improvement also are discussed.

Medical subject headings