EAAR: Efficient and Accurate Action Recognition model with enhanced spatio-temporal perception.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40652907.
- Also identified by DOI 10.1016/j.neunet.2025.107737.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Video understanding has a wide range of applications in many domains, such as security, transportation, and entertainment, with action recognition serving as a foundation to facilitate efforts such as label generation and gesture capture. Its optimization strategy revolves around efficient and accurate video data processing, improving user experience and operational efficiency. We propose an Efficient and Accurate Action Recognition (EAAR) model within a 2D CNN to achieve these goals. EAAR enhances the simultaneous sampling of temporal and spatial information, combining global and local features to reduce the interference of irrelevant information and obtain a characteristic representation of the action. It consists of two key components: the Spatial-Temporal Fusion Module (STFM) and the Spatial-Temporal Optimized Module (STOM). STFM maintains the convolutional layer's computational efficiency and memory usage, enabling simultaneous spatial and temporal sensing and enhancing the spatial distribution of the action. STOM then performs global and local feature steering, optimizing the representation of spatial and temporal features on action characteristics with linearized computational complexity, thereby improving recognition accuracy. The experimental results demonstrate the effectiveness of the proposed method in utilizing spatio-temporal features for 97.7% in action recognition and 98% in violence detection.
Medical subject headings
- Neural Networks, Computer
- Pattern Recognition, Automated
- Space Perception
- Time Perception