Pichia-CLM: A language model-based codon optimization pipeline for <i>Komagataella phaffii</i>.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41701818.
- Also identified by DOI 10.1073/pnas.2522052123 and PMC identifier 12933070.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
The preference in synonymous codon usage-the so-called codon usage bias (CUB)-is governed by several factors such as the host organism, context and function of the gene, and the position of the codon within the gene itself. We demonstrated that this mapping can be learned from the host's genome using language models and subsequently applied for codon optimization of heterologous proteins expressed by the host. This pipeline called Pichia-Codon language model (Pichia-CLM) was applied to the industrial host organism, Komagataella phaffii. With this approach, production of heterologous proteins was enhanced up to threefold compared to their native sequences. Furthermore, Pichia-CLM consistently yielded constructs with enhanced productivity for proteins of varied complexity, compared to commercially available tools. Finally, we showed that Pichia-CLM generates sequences resembling the properties of codon usage found in the host's intrinsic host cell proteins and learned features such as avoiding negative cis-regulatory and repeat elements based on patterns in the genome data. These results show the potential of language models to unbiasedly learn patterns and design robust sequences for improved protein production.
Medical subject headings
- Saccharomycetales
- Codon Usage
- Codon
- Pichia
- Models, Genetic