EPSOL: sequence-based protein solubility prediction using multidimensional embedding.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 34145885.
- Also identified by DOI 10.1093/bioinformatics/btab463.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The heterologous expression of recombinant protein requires host cells, such as Escherichiacoli, and the solubility of protein greatly affects the protein yield. A novel and highly accurate solubility predictor that concurrently improves the production yield and minimizes production cost, and that forecasts protein solubility in an E.coli expression system before the actual experimental work is highly sought. In this article, EPSOL, a novel deep learning architecture for the prediction of protein solubility in an E.coli expression system, which automatically obtains comprehensive protein feature representations using multidimensional embedding, is presented. EPSOL outperformed all existing sequence-based solubility predictors and achieved 0.79 in accuracy and 0.58 in Matthew's correlation coefficient. The higher performance of EPSOL permits large-scale screening for sequence variants with enhanced manufacturability and predicts the solubility of new recombinant proteins in an E.coli expression system with greater reliability. EPSOL's best model and results can be downloaded from GitHub (https://github.com/LiangYu-Xidian/EPSOL). Supplementary data are available at Bioinformatics online.
Medical subject headings
- Proteins
- Escherichia coli