Hierarchical memory-based deep reinforcement learning in simulated survival environments.

Cheng, Yuhu; Zhang, Yuequn; Chen, C L Philip; Wang, Xuesong · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

In complex decision-making tasks, agents often suffer from memory interference, inefficient exploration under sparse rewards, and limited adaptability in dynamic environments. While existing deep reinforcement learning (DRL) methods often suffer from performance limitations when handling long-horizon tasks due to memory interference, insufficient exploration, and poor adaptability. Inspired by multi-timescale memory mechanisms in neuroscience, this article proposes a hierarchical memory-based DRL (HM-DRL) architecture that integrates three complementary memory layers: a perceptual memory layer for real-time detection of environmental changes, an episodic memory layer for structured event sequence storage, and an abstract memory layer that constructs causal graph and supports backward reasoning. Furthermore, HM-DRL incorporates a dynamic gating mechanism and a compound reward function to facilitate effective integration of the policy network and memory outputs, thereby enhancing policy optimization and crisis response. Within the designed open-ended multi-task simulated survival environment, the HM-DRL architecture demonstrates substantial improvements in mitigating memory interference during long-horizon tasks, enhancing adaptability to sudden crises and dynamic environments, and improving learning efficiency under sparse reward scenarios. This article offers an effective paradigm for addressing core memory management challenges in long-horizon decision-making and establishes a foundation for developing agents with causal reasoning capabilities.