ORAL: Adaptive Gap Increasing for Advantage Learning via Occam's Razor Principle.

Zhang, Zhe; Zhou, Yongle; Long, Yuyang; Zhang, Jia; Weng, Juanjuan; Li, Zhetao; Gan, Yaozhong; Tan, Xiaoyang · IEEE Trans Neural Netw Learn Syst · 2026

Where this comes from

Abstract

Benefiting from the gap increasing between the optimal action and its competitors, the advantage learning (AL) operator is more robust to estimation errors in the approximated $Q$ -functions than the Bellman optimality operator in reinforcement learning (RL). However, our analysis reveals that its robustness and larger action gaps come at the cost of a worse performance loss bound, leading to slower convergence of value functions. To address this issue, we present a novel method, named Occam's Razor-based AL (ORAL), which follows Occam's Razor principle and takes the necessity into consideration when increasing the action gap. Specifically, our ORAL can adaptively increase the action gap for different state-action pairs, depending on the proximity of their $Q$ values to the optimal ones. We first propose a naive implementation of ORAL, employing a nonsmooth clipping function to realize the above idea, and then introduce a smooth version of ORAL aimed at achieving more stable learning. Furthermore, our methods can be easily plugged into other AL-based operators and extended to more complex continuous-control tasks. Theoretical analysis supports the feasibility of our approaches, demonstrating their ability to balance the gap increasing with fast convergence. Empirical results further validate its effectiveness, showing significant performance improvements across multiple benchmarks.