Exploring homology detection via k-means clustering of proteins embedded with a large language model.

Minotto, Thomas; Claessens, Antoine; Otto, Thomas D · Bioinformatics · 2025

basic_science · Level V

Where this comes from

Abstract

Inferring protein homology from sequence information is essential for understanding species evolution and enabling functional annotation transfer. Besides similarity-based methods, several machine learning approaches have been developed using various ways of representing protein data. Here, we represent proteins with a biologically oriented large language model and apply k-means clustering to the embedded data to extract homology relationships. Although our approach lacks the sensitivity of other tools, we obtain better precision for the detection of n:m orthologs. Furthermore, we successfully reconstruct full orthologous groups from scratch, highlighting the growing potential of using large language models in combination with clustering algorithms for the analysis of protein data. Datasets are available on OrthoMCL-DB as indicated in the Methods. Source code is available on GitHub at https://github.com/ThomasGTHB/OrthoLM and Zenodo at https://doi.org/10.5281/zenodo.16640170.

Medical subject headings