A New Accelerated Off-Policy Stochastic Preconditioned TD(0) Algorithm.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40408195.
- Also identified by DOI 10.1109/TPAMI.2025.3572807.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In this article, we consider policy evaluation in off-policy reinforcement learning and propose a novel procedure (Stochastic Preconditioned Temporal Difference (SPTD)) that achieves the optimal convergence rate under linear function approximation. The procedure has a linear computational complexity of the dimension of the feature space in each iteration. Under Markovian sampling, we establish finite-sample rates when the target policy can be different from the behavior policy for data generation. Our procedure is the first algorithm for the off-policy policy evaluation that has the optimal rate $\mathcal {O}(1/t)$O(1/t) under the mean square error. We also provide the first result on the asymptotic distribution and give the nearly optimal step size $\alpha _{t} = \mathcal {O}(t^{-2/3})$αt=O(t-2/3). The numerical performance of the procedure is studied in both on-policy and off-policy settings. Extensive numerical experiments demonstrate that our procedure uniformly outperforms existing methods.