Exploring homology detection via k-means clustering of proteins embedded with a large language model.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40857396.
- Also identified by DOI 10.1093/bioinformatics/btaf472 and PMC identifier 12517335.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Inferring protein homology from sequence information is essential for understanding species evolution and enabling functional annotation transfer. Besides similarity-based methods, several machine learning approaches have been developed using various ways of representing protein data. Here, we represent proteins with a biologically oriented large language model and apply k-means clustering to the embedded data to extract homology relationships. Although our approach lacks the sensitivity of other tools, we obtain better precision for the detection of n:m orthologs. Furthermore, we successfully reconstruct full orthologous groups from scratch, highlighting the growing potential of using large language models in combination with clustering algorithms for the analysis of protein data. Datasets are available on OrthoMCL-DB as indicated in the Methods. Source code is available on GitHub at https://github.com/ThomasGTHB/OrthoLM and Zenodo at https://doi.org/10.5281/zenodo.16640170.
Medical subject headings
- Proteins
- Sequence Analysis, Protein
- Computational Biology
- Sequence Homology, Amino Acid