BPA: a BERT-based priority annotation strategy for assessing the rationality of aquatic algal protein sequences.

Huang, Rui-Hua; Liang, Jun-Ze; Sun, Zheng-Hua; Chen, Xiang-Wu; Wei, Mei-Hua; Zeng, Yu-Jie; Fan, Zi-Hong; He, Qing-Yu et al. · Brief Bioinform · 2025

basic_science · Level V

Where this comes from

Abstract

Database searching remains the main approach for mass spectrometry-based proteomics, where protein identification fundamentally requires prior inclusion in the reference database. For aquatic algal species lacking annotated genomes, six-frame translation of species-specific transcriptomes has emerged as a prevalent method. However, this approach results in databases that encompass all potential translation products, substantially increasing the database size and search space. Here, we introduce BERT-based Protein Annotation (BPA), a deep learning strategy that combines a pretrained BERT model for contextual patterns, Pseudo Amino Acid Composition for physicochemical properties, and InterProScan for functional domain prediction, to optimize reference proteome construction. These features are integrated by using a Random Forest classifier to generate dynamic Sequence Reliability Scores, enabling adaptive filtering thresholds tailored to diverse experimental designs. Based on the validation across three distinct test species, this study demonstrates a robust performance of BPA with sustained high classification accuracy (AUC > 0.95). In the application to Karenia mikimotoi, BPA achieved 90% proteome compression while maintaining 40% identification coverage, effectively resolving the peptide ambiguity from redundant translations. This framework provides a scalable and efficient solution for constructing and optimizing reference libraries, facilitating proteomic research in aquatic algae and other genomically understudied species. Source code and executables are available at (https://github.com/huangruihua/BPA.git).

Medical subject headings