CIMB-MVQA: Causal intervention on modality-specific biases for medical visual question answering.

Liu, Bing; Liu, Lijun; Ding, Jiaman; Yang, Xiaobing; Peng, Wei; Liu, Li · Med Image Anal · 2026

basic_science · Level V

Where this comes from

Abstract

Medical Visual Question Answering (Med-VQA) systems frequently rely on spurious visual and language cues produced by dataset biases and structural con-founders, which undermines robustness and real-world generalization. To alleviate spurious cue reliance attributable to particular confounders, we propose CIMB-MVQA, a framework for Causal Intervention on Modality-specific Biases, which suppresses cross-modal bias by explicitly modeling and adjusting for confounding factors. For unobservable visual confounders, we introduce a front-door adjustment pipeline combining contrastive representation learning, feature disentanglement, and dual semantic masking to eliminate co-occurring but non-causal visual patterns. For observable linguistic confounders, we apply a back-door adjustment strategy using a global language bias dictionary to detect spurious signals. A vision-guided pseudo-token injection mechanism is further designed to embed critical visual cues into the language stream, reducing language dominance and aligning causal semantics across modalities. This is followed by a causal graph reasoning module that explicitly intervenes in bias-inducing paths. Experiments on multiple Med-VQA benchmarks demonstrate that CIMB-MVQA significantly improves answer accuracy and causal interpretability. Additionally, on the curated imbalanced VQA-RAD* and a suite of controlled-shift datasets, confounder-level experiments consistently show robust causal generalization under realistic bias conditions. The source code is publicly available at https://github.com/cloneiq/CIMB-MVQA.

Medical subject headings