Cross-Modality Causal-Aware Hierarchical Representation Learning for Domain Generalized Object Detection.

Liu, Yajing; Liu, Zhiyuan; Liu, Xiyao; Fan, Baojie; Tian, Jiandong · IEEE Trans Pattern Anal Mach Intell · 2026

Where this comes from

Abstract

Domain Generalized Object Detection (DGOD) addresses the critical challenge of detecting objects across diverse unseen visual domains. Recent advances in vision-language models (VLMs) have shown promising zero-shot generalization capabilities that benefit DGOD. However, existing VLM-based DGOD methods primarily leverage VLMs for data augmentation, overlooking the rich generalizable knowledge it contains. Besides, VLM-based methods developed for other domain generalization (DG) tasks suffer from modality gap and representation bias due to their correlation-driven cross-modal interaction paradigms, which severely limit the performance to DGOD. To bridge this research gap and advance the causal-driven VLM-based DG, we develop a Causal-aware Hierarchical Representation Graph that reformulates the VLM-based DG problem into a hierarchical causal representation learning framework. Our framework incorporates a Generalizable Knowledge Transfer module to refine and inherit transferable scene-object features from the VLM's visual encoder, and a Causal Prototype Learning module that employs general text embeddings as causal interventions to guide the construction of a mediator visual causal prototype space, which inherits the generalization and category relational representation ability from text without representation bias. Furthermore, we introduce a Prototypical Cross-attention Classifier that eliminates modality gap by integrating object features with learned causal prototypes for text-free classification, which also directs visual features to approach the mediator causal space, enabling the learning of causal visual features that are invariant to domain-specific confounders. Our causal-driven framework transfers causal invariance from text to visual modality and provides a vision-friendly perspective for leveraging VLMs to solve vision-centric tasks. Extensive experiments on five benchmarks including Diverse Weather, Corruption, Real-to-Artistic, Cross Camera and Sim-to-Real demonstrate that our method achieves superior generalization performance.