Clustering the protein universe of life using DIAMOND DeepClust.
Where this comes from
- Record sourced from PubMed, PMID 41876643.
- Also identified by DOI 10.1038/s41592-026-03030-z and PMC identifier 13076203.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Relating billions of proteins across the tree of life remains a challenging task for comparative biosphere genomics and artificial intelligence-driven structure prediction. Here we present DIAMOND DeepClust, a cascaded, ultra-fast clustering method enabling planetary-scale organization of protein space, scaling to trillions of sequences while retaining sensitivity at low identity. Aggregating 19 billion biosphere proteins into 544 million nonsingleton clusters, we show that using our DeepClust database, available for download, can enhance structure prediction with AlphaFold2.
Medical subject headings
- Proteins
- Computational Biology