Emotion-Aware multimodal deepfake detection.

Zhang, Teng; Li, Gen; Xiao, Yanhui; Tian, Huawei; Cao, Yun · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

With the continuous advancement of Deepfake techniques, traditional unimodal detection methods struggle to address the challenges posed by multimodal manipulations. Most existing approaches rely on large-scale training data, which limits their generalization to unseen identities or different manipulation types in few-shot settings. In this paper, we propose an emotion-aware multimodal Deepfake detection method that exploits emotion signals for forgery detection. Specifically, we design an emotion embedding extractor (Emoencoder) to capture emotion representations within modalities. Then, we employ Emotion-Aware Contrastive Learning and Cross-Modal Contrastive Learning to capture cross-modal inconsistencies and enhance modality feature extraction. Furthermore, we propose a Text-Guided Semantic Fusion module, where the text modality serves as a semantic anchor to guide audio-visual feature interactions for multimodal feature fusion. To validate our approach under data-limited conditions and unseen identities, we employ a cross-identity few-shot training strategy on benchmark datasets. Experimental results demonstrate that our method outperforms SOTAs and demonstrates superior generalization to both unseen identities and manipulation types.

Medical subject headings