Exploring structural diversity across the protein universe with The Encyclopedia of Domains.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 39480926.
- Also identified by DOI 10.1126/science.adq4946 and PMC identifier 7618865.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
The AlphaFold Protein Structure Database (AFDB) contains more than 214 million predicted protein structures composed of domains, which are independently folding units found in multiple structural and functional contexts. Identifying domains can enable many functional and evolutionary analyses but has remained challenging because of the sheer scale of the data. Using deep learning methods, we have detected and classified every domain in the AFDB, producing The Encyclopedia of Domains. We detected nearly 365 million domains, over 100 million more than can be found by sequence methods, covering more than 1 million taxa. Reassuringly, 77% of the nonredundant domains are similar to known superfamilies, greatly expanding representation of their domain space. We uncovered more than 10,000 new structural interactions between superfamilies and thousands of new folds across the fold space continuum.
Medical subject headings
- Databases, Protein
- Protein Domains
- Protein Folding
- Proteins
- Deep Learning