Bidirectional transition consistency between multi-domain observations for visual reinforcement learning generalization.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41218402.
- Also identified by DOI 10.1016/j.neunet.2025.108265.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Visual reinforcement learning has proven effective in addressing control tasks with high-dimensional image observations. However, obtaining generalizable representations for policy learning under various visual interferences remains a significant challenge. Inspired by human cognitive processes in unfamiliar scenarios, we propose the Multi-Domain Bidirectional Transition (MDBT) model. Unlike prior approaches that either enforce visual consistency or rely solely on precise model-based transitions, MDBT explicitly incorporates multi-domain observations with visual perturbations and kernel regions, while emphasizing task-related dynamics. This design enables MDBT to remove noise interference and preserve task relevance, resulting in more robust and transferable representations.MDBT consists of three key components. Initially, we utilize the Data Transformation module to diversify the original observations, thereby obtaining multi-domain observations with various degrees of visual interference. Subsequently, we employ a Bidirectional Transition module to predict both forward and backward environmental transitions, extracting the task-relevant representations. Finally, we impose a Consistency Target to constrain the coherent multi-domain transition predictions, ensuring that the representations remove noise interference while retaining task relevance. Extensive experiments demonstrate that MDBT achieves consistent state-of-the-art performance: it outperforms prior approaches by an average of 2.3% in success rate on the "Video Hard" setting of the DeepMind Control Suite, improves average generalization by up to 124.1% on the "Reach" task and 73.5% on the "Peg in Box" task of robotic manipulation benchmarks, and enhances average generalization returns by up to 23.2% and average driving distance by up to 16.9% under severe weather and lighting perturbations in CARLA. These results highlight the effectiveness of MDBT in learning robust and transferable representations for visual reinforcement learning.
Medical subject headings
- Reinforcement, Psychology
- Generalization, Psychological
- Neural Networks, Computer
- Visual Perception