Toward global interpretability of neural networks via Boolean transformation.

Tang, Yiping; Shao, Rui; Dai, Ruifen; Wang, Fang · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Interpretability remains a central challenge in the deployment of deep neural networks, particularly in safety-critical and decision-sensitive fields. This work proposes a unified framework for post-hoc global interpretability by transforming general neural network architectures-including the Residual Network and Transformer-into equivalent decision diagrams over real-valued inputs and multi-class outputs. These decision diagrams provide a transparent, structured view of the neural network's overall behavior, where each path encodes a tractable and interpretable decision rule. We identify a counterintuitive yet effective modification in the node merging process during diagram construction, which leads to faster entropy reduction and smaller equivalent intervals, thereby significantly reducing the diagram size while maintaining equivalence with the original network. The resulting representations not only support the exploration of logical properties, such as decision boundary tracing, equivalence checking, robustness analysis, and model counting, but also serve as globally interpretable surrogates for the original neural networks. Experiments validate the effectiveness and scalability of the proposed methods, highlighting their potential for reliable neural network analysis and verification.