A Survey on Human-Centric Voice-Face Multimodal Learning.

Chen, Wuyang; Xu, Kele; Song, Qiya; Sun, Yanjie; Liu, Xinwang; Li, Xingyi; Dou, Yong; Wang, Huaimin · IEEE Trans Neural Netw Learn Syst · 2026

review · Level V

Where this comes from

Abstract

The integration of voice-encompassing both speech and nonverbal acoustics-with facial data in human-centric voice-face multimodal learning has emerged as a critical paradigm for understanding human-centered behavioral patterns. Unlike general audio-visual learning, it is deeply rooted in biometric and neurocognitive principles. However, existing studies often remain task-specific, lacking a holistic perspective that recognizes their interdependencies and neglecting the unique data, model, and task properties inherent to human-centric learning. This survey provides a systematic overview of voice-face multimodal learning by categorizing research into five key areas: 1)biometric and neurocognitive foundations, establishing the theoretical underpinnings of voice-face perception and providing insights for model design; 2)task evolution and correlation, revealing hidden dependencies between various tasks through an evolutionary trajectory analysis; 3)human-centric representation learning, designed for multimodal data with human-specific characteristics; and 4)dataset taxonomy, structuring datasets based on annotation granularity, desired properties, and task compatibility; and 5)demographic bias and task-specific edge cases, analyzing fairness, robustness inherent to human-centric voice-face learning. We review both representative past approaches and the latest advancements through several typical downstream tasks, highlight open challenges, and outline future directions, aiming to unify fragmented research and inspire advancements in this field. More details are available online (https://github.com/audio-visual/Survey-on-voice-face-multimodal-learning).