A New Accelerated Off-Policy Stochastic Preconditioned TD(0) Algorithm.

Liu, Weidong; Ma, Jiahua; Mao, Xiaojun; Tang, Kejie · IEEE Trans Pattern Anal Mach Intell · 2025

basic_science · Level V

Where this comes from

Abstract

In this article, we consider policy evaluation in off-policy reinforcement learning and propose a novel procedure (Stochastic Preconditioned Temporal Difference (SPTD)) that achieves the optimal convergence rate under linear function approximation. The procedure has a linear computational complexity of the dimension of the feature space in each iteration. Under Markovian sampling, we establish finite-sample rates when the target policy can be different from the behavior policy for data generation. Our procedure is the first algorithm for the off-policy policy evaluation that has the optimal rate $\mathcal {O}(1/t)$O(1/t) under the mean square error. We also provide the first result on the asymptotic distribution and give the nearly optimal step size $\alpha _{t} = \mathcal {O}(t^{-2/3})$αt=O(t-2/3). The numerical performance of the procedure is studied in both on-policy and off-policy settings. Extensive numerical experiments demonstrate that our procedure uniformly outperforms existing methods.