A multimodal transformer-based visual question answering method integrating local and global information.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40601713.
- Also identified by DOI 10.1371/journal.pone.0324757 and PMC identifier 12220993.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Addressing the limitations in current visual question answering (VQA) models face limitations in multimodal feature fusion capabilities and often lack adequate consideration of local information, this study proposes a multimodal Transformer VQA network based on local and global information integration (LGMTNet). LGMTNet employs attention on local features within the context of global features, enabling it to capture both broad and detailed image information simultaneously, constructing a deep encoder-decoder module that directs image feature attention based on the question context, thereby enhancing visual-language feature fusion. A multimodal representation module is then designed to focus on essential question terms, reducing linguistic noise and extracting multimodal features. Finally, a feature aggregation module concatenates multimodal and question features to deepen question comprehension. Experimental results demonstrate that LGMTNet effectively focuses on local image features, integrates multimodal knowledge, and enhances feature fusion capabilities.
Medical subject headings
- Visual Perception