Pan-microalgal dark proteome mapping via interpretable deep learning and synthetic chimeras.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41328162.
- Also identified by DOI 10.1016/j.patter.2025.101373 and PMC identifier 12664985.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Microalgal genomes contain a vast "dark proteome"-sequences lacking detectable homology that evade conventional classification tools. We developed LA<sup>4</sup>SR (language modeling with AI for algal amino acid sequence representation), a framework using transformer- and state-space models to classify translated ORFeomes across ten algal phyla. Training on ∼77 million sequences, LA<sup>4</sup>SR achieves near-complete recall, accelerates classification by ∼10,701× relative to BLASTP<sup>+</sup>, and generalizes robustly to unseen sequences using less than 2% of available data. Models trained on synthetic, chimeric (terminal information [TI]-free) sequences maintained high accuracy, demonstrating that internal sequence features alone can drive robust classification. Inference speed and scalability were further enhanced under TI-free settings, supporting rapid annotation of large proteomic datasets. Custom explainability tools revealed interpretable amino acid patterns linked to evolutionary and biophysical features. Designed for accessibility across disciplines, LA<sup>4</sup>SR integrates biological context and computational innovation in parallel, enabling both biologists and data scientists to interrogate the microbial dark proteome.