Mechanistic Interpretability of Fine-Tuned Protein Language Models for Nanobody Thermostability Prediction.

Murakami, Taihei; Hashidate, Yuki; Matsunaga, Yasuhiro · Bioinformatics · 2026

basic_science · Level V

Where this comes from

Abstract

While Protein Language Models (PLMs) fine-tuned on biophysical data achieve high predictive accuracy, the physical principles underlying their predictions remain obscure. Deciphering these representations offers a unique opportunity to not only interpret model decisions but also to discover novel biophysical insights governing protein properties. Here, we present a framework using Sparse Autoencoders (SAEs) to extract mechanistic knowledge from PLMs fine-tuned for nanobody thermostability. We fine-tuned the ESM-2 model on the nanobody thermostability dataset, achieving superior performance compared to significantly larger state-of-the-art models. SAE analysis successfully decomposed the model's dense embeddings into sparse, interpretable features without loss of predictive accuracy. We characterized these features through both global and local analyses. Global analysis provided an aggregate map of position-dependent feature contributions, whereas local analysis identified specific residue-level patterns, including known determinants such as the VHH-tetrad and critical disulfide bonds, as well as candidate stabilizing residues. Free Energy Perturbation calculations supported the structural plausibility of selected residue-level hypotheses. These results show that SAE-based interpretation can generate testable, structurally grounded hypotheses for rational protein engineering. The data and source code of the proposed method are available at GitHub (https://github.com/matsunagalab/paper_nanobody-thermostability-sae) and Zenodo (DOI: 10.5281/zenodo.18012027). Supplementary data are available at Bioinformatics online.