Language-Driven Spatial-Semantic Cross-Attention for Face Attribute Recognition With Limited Labeled Data.

Kim, Young-Eun; Bak, Gyeong-Min; Lee, Seong-Whan · IEEE Trans Neural Netw Learn Syst · 2025

basic_science · Level V

Where this comes from

Abstract

Recent advances in deep learning have demonstrated excellent results for face attribute recognition (FAR), which is generally trained with large-scale labeled data. Despite the significant progress in this field, most existing works mainly rely on large-scale labeled data, which is not practical in many real-world FAR applications. Numerous studies have been conducted to address this problem, but they require either large external face datasets or complex auxiliary tasks for pretraining the backbone network. In this article, we propose a new method named language-driven spatial-semantic cross-attention (LSA) that does not require any pretraining steps with additional datasets or auxiliary tasks. Driven by the impressive outcomes of recent computer vision studies using language models, we harness language-based relational information to enhance attribute recognition. The core of LSA is to combine and balance the learned scaled-dot product attention with the attention constructed based on language-driven knowledge. To this end, we propose a correlation dictionary, obtained with the similarity between text embeddings of facial attributes and facial regions to represent relationships. The correlation dictionary then creates a cross-attention form and is combined into the cross-attention with balancing parameters. Thus, we can compensate for the lack of data information by providing prior knowledge directly to the network. Extensive experiments demonstrate that our method surpasses state-of-the-art techniques, achieving an average improvement of 0.29% on the CelebA dataset and 0.39% on the LFWA dataset with limited labeling data, even without additional dataset training.