Toward global interpretability of neural networks via Boolean transformation.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42526151.
- Also identified by DOI 10.1016/j.neunet.2026.109419.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Interpretability remains a central challenge in the deployment of deep neural networks, particularly in safety-critical and decision-sensitive fields. This work proposes a unified framework for post-hoc global interpretability by transforming general neural network architectures-including the Residual Network and Transformer-into equivalent decision diagrams over real-valued inputs and multi-class outputs. These decision diagrams provide a transparent, structured view of the neural network's overall behavior, where each path encodes a tractable and interpretable decision rule. We identify a counterintuitive yet effective modification in the node merging process during diagram construction, which leads to faster entropy reduction and smaller equivalent intervals, thereby significantly reducing the diagram size while maintaining equivalence with the original network. The resulting representations not only support the exploration of logical properties, such as decision boundary tracing, equivalence checking, robustness analysis, and model counting, but also serve as globally interpretable surrogates for the original neural networks. Experiments validate the effectiveness and scalability of the proposed methods, highlighting their potential for reliable neural network analysis and verification.