The more, the merrier: Detecting categorical emotions from texts with cross-modal insights.

Min, Changrong; Wang, Aimin; Li, Ximing · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Categorical Emotion Detection (CED) task aims to identify the emotion expressed in a given text. While Prompt Tuning (PT) has been applied to the CED, existing detectors struggle to design task-efficient prompts. Given that human perception is cross-modal, we propose a Visual Emotion Prefix-guided Emotion Detector (VisPor) that exploits visual information from emotion resources to heuristically construct prompts that better adapt to CED. In particular, to mitigate the severe modality gap, VisPor aligns abstract textual descriptions of emotions with concrete emotional images within a shared space. Then, by taking such visual-enriched text embeddings as a prefix, the visual emotion information can smoothly and effectively serve VisPor to capture nuanced emotion features from texts. We conducted extensive experiments on three CED benchmarks and the results show that VisPor significantly outperforms the existing methods.

Medical subject headings