Cross-Modality Causal-Aware Hierarchical Representation Learning for Domain Generalized Object Detection.
Where this comes from
- Record sourced from PubMed, PMID 42743023.
- Also identified by DOI 10.1109/TPAMI.2026.3733780.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Domain Generalized Object Detection (DGOD) addresses the critical challenge of detecting objects across diverse unseen visual domains. Recent advances in vision-language models (VLMs) have shown promising zero-shot generalization capabilities that benefit DGOD. However, existing VLM-based DGOD methods primarily leverage VLMs for data augmentation, overlooking the rich generalizable knowledge it contains. Besides, VLM-based methods developed for other domain generalization (DG) tasks suffer from modality gap and representation bias due to their correlation-driven cross-modal interaction paradigms, which severely limit the performance to DGOD. To bridge this research gap and advance the causal-driven VLM-based DG, we develop a Causal-aware Hierarchical Representation Graph that reformulates the VLM-based DG problem into a hierarchical causal representation learning framework. Our framework incorporates a Generalizable Knowledge Transfer module to refine and inherit transferable scene-object features from the VLM's visual encoder, and a Causal Prototype Learning module that employs general text embeddings as causal interventions to guide the construction of a mediator visual causal prototype space, which inherits the generalization and category relational representation ability from text without representation bias. Furthermore, we introduce a Prototypical Cross-attention Classifier that eliminates modality gap by integrating object features with learned causal prototypes for text-free classification, which also directs visual features to approach the mediator causal space, enabling the learning of causal visual features that are invariant to domain-specific confounders. Our causal-driven framework transfers causal invariance from text to visual modality and provides a vision-friendly perspective for leveraging VLMs to solve vision-centric tasks. Extensive experiments on five benchmarks including Diverse Weather, Corruption, Real-to-Artistic, Cross Camera and Sim-to-Real demonstrate that our method achieves superior generalization performance.