Multi-modal feature alignment networks for multi-label image classification.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41579821.
- Also identified by DOI 10.1016/j.neunet.2026.108629.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Multi-label image classification is a classification task that assigns labels to multiple objects in an input image. Recent research ideas mainly focus on solving the semantic consistency of visual features and label features. However, since images contain complex scene content, the features captured by visual feature extraction networks based on grid or sequence representation may introduce redundant information or lack continuity when identifying irregular objects. In order to fully mine the visual information of complex objects in images and enhance the inter-modal interaction of images and labels, we introduce a flexible graph structure to explore the internal information of objects and design a multi-modal feature alignment (MMFA) network for multi-label image classification. To enhance the context awareness and semantic association of different patch regions, we propose a semantic-augmented interaction module that combines two kinds of visual semantic information with label embeddings for interactive learning. Finally, we refine the dependence between local intrinsic information and overall semantics by redefining semantic queries through semantically enhanced visual spatial features and graph aggregation features. Experiments on three large-scale public datasets: Microsoft COCO, Pascal VOC 2007 and NUS-WIDE demonstrate the effectiveness of our proposed MMFA and achieve state-of-the-art performance.
Medical subject headings
- Neural Networks, Computer
- Image Processing, Computer-Assisted
- Pattern Recognition, Automated