An efficient and low-latency attention model for event denoising.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41722301.
- Also identified by DOI 10.1016/j.neunet.2026.108729.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Event-based cameras are bio-inspired vision sensors that capture brightness changes at each pixel independently, generating sparse events with ultra-low latency. Unlike frame-based cameras, they provide highly efficient, temporally precise scene information, making them ideal for real-time processing. However, the output data from event-based cameras is often contaminated by noise, particularly Background Activity (BA) noise, which degrades data quality and affects downstream tasks. In this paper, we introduce StatFormer, a novel event denoising method designed to enhance temporal feature modeling while maintaining computational efficiency. Specifically, StatFormer incorporates statistical information into the event embedding process, allowing temporal dynamics to be directly modeled at the input level. To further enhance denoising performance, we adopt a two-stage transformer architecture focused on local attention. In the first stage, local attention is extracted within short sub-sequences to capture fine-grained spatial-temporal dependencies. In the second stage, local attention is strengthened over the entire event sequence. Compared to previous attention-based model, our approach significantly reduces both model size and inference time. Experimental results demonstrate that StatFormer achieves a 13.41 × speedup in inference, with an average processing time of just 1.38 μs per event, while delivering best denoising performance across several publicly available datasets.
Medical subject headings
- Attention
- Neural Networks, Computer