DECHRL: Empowerment-driven delay-aware causal hierarchical reinforcement learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42628460.
- Also identified by DOI 10.1016/j.neunet.2026.109507.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Many real-world tasks involve delayed effects, where the outcomes of actions emerge after varying time lags. Existing delay-aware reinforcement learning methods often rely on state augmentation, prior knowledge of delay distributions, or access to non-delayed data-limiting their generalization. Hierarchical reinforcement learning, by contrast, inherently offers advantages in handling delays due to its hierarchical structure, yet existing methods are restricted to fixed delays. To address these limitations, we propose Delay-Empowered Causal Hierarchical Reinforcement Learning (DECHRL). DECHRL consists of two modules: causal delay distribution modeling and delay-aware empowerment-driven hierarchical policy training. The former learns Granger-causal relationships from delayed interactions, while the latter uses this information to construct and optimize hierarchical policies. By guiding exploration toward controllable and informative states, DECHRL improves exploration efficiency and decision stability under temporal uncertainty. We evaluate DECHRL on modified 2D-Minecraft and MiniGrid environments with constructed stochastic delays, and further extend the evaluation to the low-dimensional continuous-state PointMaze task. Experimental results show that DECHRL effectively models causal delay distributions and significantly outperforms baselines under temporal uncertainty. The additional PointMaze results further demonstrate the applicability of the proposed delay-aware causal modeling mechanism to low-dimensional continuous navigation tasks with stochastic delayed transitions.