Pseudo-distribution elite critics: Enhancing accuracy in reinforcement learning value estimation.

Zhang, Yujia; Li, Lin; Wei, Wei; Wu, Jianguo; Liang, Jiye · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Reinforcement learning has succeeded significantly in developing intelligent agents capable of adeptly navigating complex environments, yet it often encounters limitations due to persistent biases in state-action value estimation. To address this challenge, we introduce the Pseudo-distribution Elite Critics (PEC), an innovative framework designed to enhance sample efficiency and effectively balance overestimation and underestimation biases in Q-value approximations. PEC revolves around the innovative concept of pseudo-distribution representation, which enriches Q-value approximations with distributional characteristics, capturing nuanced variations in Q-values without increasing the number of critics, leading to more refined and precise estimations. This framework is further enhanced by two critical components: an uncertainty measurement that accurately identifies the most reliable critic for Temporal Difference target computation and a trimmed mean technique that adeptly balances optimistic and pessimistic biases in Temporal Difference target. Empirical studies across various benchmark scenarios validate the statistical significance of PEC and its superior performance over existing methodologies.

Medical subject headings