Advancing Zero-Shot Adversarial Robust Fairness: Theory and Algorithm.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42752408.
- Also identified by DOI 10.1109/TPAMI.2026.3735010.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Vision-language models (VLMs) exhibit strong zero-shot generalization across downstream tasks. However, they remain highly vulnerable to adversarial perturbations that are nearly imperceptible to humans. While recent studies have proposed various defense techniques to enhance the average adversarial robustness of VLMs, little attention has been paid to the class-wise disparities in robustness, a critical issue that we refer to as zero-shot robust fairness. In this work, we conduct a comprehensive analysis of robust fairness in CLIP-based models from both theoretical and empirical perspectives, and identify that the decreased inter-class separation and increased intra-class dispersion are key factors leading to inconsistencies of robustness across classes. To conquer the above two challenges, the von Mises-Fisher (vMF) distribution that resorts to characterize the concentration of data in high-dimensional space is deployed. In particular, we employ vMF to quantify the relationships of features, and propose two tailored loss functions to mitigate class-wise disparities by encouraging inter-class separation and enhancing intra-class compactness of VLMs. The proposed method is modular and can be seamlessly plugged and played in existing adversarial fine-tuning pipelines. Extensive experiments on 15 downstream benchmark datasets demonstrate that our method improves both the adversarial robustness and fairness of VLMs in zero-shot settings, thereby showcasing its generality and effectiveness.