Hierarchical memory-based deep reinforcement learning in simulated survival environments.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42013632.
- Also identified by DOI 10.1016/j.neunet.2026.108987.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In complex decision-making tasks, agents often suffer from memory interference, inefficient exploration under sparse rewards, and limited adaptability in dynamic environments. While existing deep reinforcement learning (DRL) methods often suffer from performance limitations when handling long-horizon tasks due to memory interference, insufficient exploration, and poor adaptability. Inspired by multi-timescale memory mechanisms in neuroscience, this article proposes a hierarchical memory-based DRL (HM-DRL) architecture that integrates three complementary memory layers: a perceptual memory layer for real-time detection of environmental changes, an episodic memory layer for structured event sequence storage, and an abstract memory layer that constructs causal graph and supports backward reasoning. Furthermore, HM-DRL incorporates a dynamic gating mechanism and a compound reward function to facilitate effective integration of the policy network and memory outputs, thereby enhancing policy optimization and crisis response. Within the designed open-ended multi-task simulated survival environment, the HM-DRL architecture demonstrates substantial improvements in mitigating memory interference during long-horizon tasks, enhancing adaptability to sudden crises and dynamic environments, and improving learning efficiency under sparse reward scenarios. This article offers an effective paradigm for addressing core memory management challenges in long-horizon decision-making and establishes a foundation for developing agents with causal reasoning capabilities.