ESTSformer: Efficient spatio-temporal spiking transformer.

Lu, Chengzhuo; Du, Huilin; Wei, Wenjie; Sun, Qian; Wang, Yuchen; Zeng, Dingyi; Chen, Wenyu; Zhang, Malu et al. · Neural Netw · 2025

basic_science · Level V

Where this comes from

Abstract

Bio-inspired Spiking Neural Networks (SNNs) have garnered significant attention for their binary, asynchronous, and event-driven computing. These characteristics make SNNs a compelling alternative to traditional Artificial Neural Networks (ANNs). As Transformers expand across AI domains, integrating them with SNNs offers a promising path to high-performance and efficient models. A critical challenge in this area is effectively leveraging SNNs' inherent spatio-temporal advantages. Unfortunately, existing studies on spiking spatio-temporal attention mechanisms exhibit certain limitations. In this paper, we first analyze the drawbacks of vanilla spatiotemporal self-attention (STSA), specifically its substantial computational and storage demands that escalate with increasing time steps. To address this, we propose an efficient spatiotemporal self-attention (ESTSA) mechanism. ESTSA divides attention heads into distinct sets for temporal and spatial information extraction, a simple yet effective division that significantly reduces both computational and storage overhead. Based on the ESTSA, we construct an efficient spiking transformer architecture, termed ESTSformer, which modifies residual connections within its encoder modules to ensure purely spike-driven computation throughout the network. Extensive experiments on both neuromorphic and static datasets demonstrate that our method achieves superior performance and efficiency compared to advanced existing works.

Medical subject headings