Visual reinforcement learning via sequential consistency preserved policy contrast from optimal transport view.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40884894.
- Also identified by DOI 10.1016/j.neunet.2025.108019.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Data inefficiency has long posed a significant challenge in the application of visual reinforcement learning methods to complex scenarios. To address this issue, recent studies have incorporated representation learning mechanisms to extract discriminative features by introducing auxiliary objectives that contrast pixel observations. However, our investigations suggest that these representations may not sufficiently capture the essential information for effective decision-making and could potentially impede policy learning. To tackle these limitations, we propose a novel methodology termed CoCo (sequential Consistency preserved policy Contrast). Unlike existing approaches, CoCo emphasizes the capture of invariant policy-based discriminative features by performing policy contrast across multiple distorted views of observations. Accordingly, we determine that there exists a certain intrinsic heterogeneity between policy and observation, since policy establishes an explicit distribution characteristic rather than a plain tensor. To this end, we propose to model the policy contrast as an optimal transport problem and further perform the alignment of policy distributions during contrastive learning. Subsequently, we introduce an inverse consistency-weighting mechanism designed to accentuate the differences between views while maintaining semantic integrity. We establish the theoretical optimality of our proposed method through an information-theoretic analysis and demonstrate its practical effectiveness via comprehensive evaluation across diverse data efficiency benchmarks, where CoCo consistently outperforms existing approaches.
Medical subject headings
- Reinforcement, Psychology
- Neural Networks, Computer