Learning Counterfactual Fair Representation Under Covariate Shift via Reflux.

Xia, Yiliang; Zhang, Xiaohang; Li, Zhengren · IEEE Trans Neural Netw Learn Syst · 2026

basic_science · Level V

Where this comes from

Abstract

Recent advances in counterfactual fairness have shifted focus from flawed group fairness metrics to ensuring individual-level fairness through counterfactual reasoning. However, most existing approaches remain limited to in-processing strategies-injecting fairness constraints into predictive models-while largely overlooking the potential of data preprocessing to mitigate inherent biases. Moreover, few methods address the critical challenge of real-world distribution shift, which can compromise the generalizability of fair models across domains. In this article, we propose the counterfactual reflux variational autoencoder (CRVAE), a novel framework for generating counterfactual samples and learning fair representations. To the best of our knowledge, this is the first work to explicitly consider counterfactual fair representation learning under covariate shift, enabling both single-domain and covariate shift prediction tasks. For fairness, we introduce a Reflux technique that enforces consistency between factual and counterfactual representations. For transferability, we incorporate a domain discriminator to align fair representations across domains. Experimental results show that our approach improves fairness with minimal performance loss and maintains generalization across domains. Furthermore, CRVAE can be flexibly combined with existing in-processing fairness methods. Future work may explore extending this framework to settings with limited causal graph knowledge.