FiNE: Fine-grained Neuron-level Model Editing for Reliable and Safe LLMs.

Yang, Xun; Pan, Haowen; Wang, Xiaozhi; Cao, Yixin; Li, Juanzi; Wang, Meng · IEEE Trans Pattern Anal Mach Intell · 2026

basic_science · Level V

Where this comes from

Abstract

The rapid advancement of Large Language Models (LLMs) has underscored the need for editing techniques that enhance both the reliability and safety of model outputs. While prior locate-then-edit methods, such as ROME, achieve factual correction via causal tracing, they often overlook the internal structure of Feed-Forward Networks (FFNs) and the relational context among edited concepts, leading to brittle or incomplete updates. To address these limitations, we propose FiNE (Fine-grained Neuron-level Editing), a unified locate-then-edit framework that performs gradient-free localization and fine-grained parameter adjustment within FFNs. FiNE introduces a theoretically grounded contribution score, which can be interpreted as an efficient variant of gradient-based methods, to identify concept-relevant neurons. It further employs a composite loss to balance editing accuracy, model consistency, and output diversity. We instantiate FiNE for two complementary applications, knowledge editing (FiNE-K) and safety editing (FiNE-S), which together demonstrate the versatility of the framework. Extensive experiments on the KnowEdit and SafeEdit benchmarks show that FiNE achieves superior precision, robustness, and efficiency with minimal disruption to the model's general behavior. These results highlight the effectiveness and scalability of fine-grained neuron editing as a unified approach to building more reliable and safer LLMs.