Visual reinforcement learning via sequential consistency preserved policy contrast from optimal transport view.

Zang, Zehua; Li, Jiangmeng; Sun, Chuxiong; Wang, Rui; Liu, Lixiang; Sun, Fuchun · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Data inefficiency has long posed a significant challenge in the application of visual reinforcement learning methods to complex scenarios. To address this issue, recent studies have incorporated representation learning mechanisms to extract discriminative features by introducing auxiliary objectives that contrast pixel observations. However, our investigations suggest that these representations may not sufficiently capture the essential information for effective decision-making and could potentially impede policy learning. To tackle these limitations, we propose a novel methodology termed CoCo (sequential Consistency preserved policy Contrast). Unlike existing approaches, CoCo emphasizes the capture of invariant policy-based discriminative features by performing policy contrast across multiple distorted views of observations. Accordingly, we determine that there exists a certain intrinsic heterogeneity between policy and observation, since policy establishes an explicit distribution characteristic rather than a plain tensor. To this end, we propose to model the policy contrast as an optimal transport problem and further perform the alignment of policy distributions during contrastive learning. Subsequently, we introduce an inverse consistency-weighting mechanism designed to accentuate the differences between views while maintaining semantic integrity. We establish the theoretical optimality of our proposed method through an information-theoretic analysis and demonstrate its practical effectiveness via comprehensive evaluation across diverse data efficiency benchmarks, where CoCo consistently outperforms existing approaches.

Medical subject headings