STFMF-Net: A Hybrid Attention-Driven Multi-View Fusion Framework for Non-Contact Anxiety Recognition Via Facial Videos.

Li, Wentong; Wu, Dan; Liu, Longxin; Li, Ye; Zhang, Qiqi; Ji, Xiaoqiang; Liu, Jikui · IEEE J Biomed Health Inform · 2026

Where this comes from

Abstract

Non-contact anxiety recognition methods based on facial videos offer a convenient and efficient approach for large-scale mental health screening. Most existing studies on anxiety recognition based on facial videos have mainly used facial expressions or remote photoplethysmography (rPPG). However, anxiety states manifest through multiple behavioral and physiological dimensions, such as fine-grained eye movement trajectories, pupil diameter, eye-blink, and rPPG signals. Therefore, single-view approaches do not fully capture information related to anxiety. To address this limitation, we propose a novel multi-view anxiety recognition framework based on facial videos. First, a multi-fusion attention net (MFA-Net) model is developed to estimate pupil diameter under complex facial video conditions, enabling precise reconstruction of pupil information. Then, we designed a multi-view fusion network (STFMF-Net) that integrates a Hybrid Fusion Module to effectively fuse the spatiotemporal and time-frequency features from eye movement signals, pupil diameter, eye-blink, head movement, and rPPG. The proposed method was evaluated on the UBFC-Phys public dataset. Experimental results indicate that our approach achieves an accuracy of 97.39% and an F1 score of 97.19%, outperforming the state-of-the-art (SOTA) by 3.06 and 2.80 percentage points, respectively. Therefore, the results highlight the effectiveness of multi-view fusion for non-contact, high-precision anxiety recognition and provide a promising technical solution for early anxiety assessment.