Connecting chemical and protein sequence space to predict biocatalytic reactions.

Paton, Alexandra E; Boiko, Daniil A; Perkins, Jonathan C; Cemalovic, Nicholas I; Reschützegger, Thiago; Gomes, Gabe; Narayan, Alison R H · Nature · 2025

basic_science · Level V

Where this comes from

Abstract

The application of biocatalysis in synthesis has the potential to offer streamlined routes towards target molecules<sup>1</sup>, tunable catalyst-controlled selectivity<sup>2</sup>, as well as processes with improved sustainability<sup>3</sup>. Despite these advantages, biocatalysis is often a high-risk strategy to implement, as identifying an enzyme capable of performing chemistry on a specific intermediate required for a synthesis can be a roadblock that requires extensive screening of enzymes and protein engineering to overcome<sup>4</sup>. Strategies for predicting which enzyme and small molecule are compatible have been hindered by the lack of well-studied biocatalytic reaction datasets<sup>5</sup>. The underexploration of connections between chemical and protein sequence space constrains navigation between these two landscapes. Here we report a two-phase effort relying on high-throughput experimentation to populate connections between productive substrate and enzyme pairs and the subsequent development of a tool, CATNIP, for predicting compatible α-ketoglutarate (α-KG)/Fe(II)-dependent enzymes for a given substrate or, conversely, for ranking potential substrates for a given α-KG/Fe(II)-dependent enzyme sequence. We anticipate that our approach can be readily expanded to further enzyme and transformation classes and will derisk the investigation and application of biocatalytic methods.

Medical subject headings