Learning representations via dynamics-based behavioral similarity for deep reinforcement learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41422620.
- Also identified by DOI 10.1016/j.neunet.2025.108468.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Learning task-relevant latent representations from visual observations is crucial for enabling effective deep reinforcement learning. Recent studies demonstrate that behavioral similarity metrics leverage reward signals and transition dynamics to group behaviorally equivalent states, thereby eliminating task-irrelevant distractions while preserving control-critical features. However, they are prone to representation collapse, primarily due to measurement inaccuracies induced by sparse reward distributions, which hinders their scalability to complex applications. To address this issue, we propose a Representation learning with Dynamics-based behavioral Similarity approach (RDS). Our key motivation lies in a dynamics-driven similarity metric that eliminates reward dependency while preserving behavioral discriminability. Specifically, we first introduce a weak behavioral similarity metric, incorporating dynamic transition distances with trainable Gaussian noise to adaptively mitigate metric degradation. Building upon this, we further exploit the latent trajectory distances unfolded through dynamic transitions to quantify the task differences of the initial states within the learned representation space, thereby extracting task-relevant features. Finally, extensive results across complex DeepMind Control, MetaWorld and Adroit manipulation tasks show that RDS outperforms competitive baselines and achieves overall improvements of 43 % and 30 % over the foundational DrQ-v2 and the state-of-the-art method, while ablation studies also confirm the effectiveness of inner components.