Value Iteration for Stochastic LQR With Convergence Guarantees.

Lai, Jing; Xiong, Junlin; Kang, Yu · IEEE Trans Neural Netw Learn Syst · 2025

basic_science · Level V

Where this comes from

Abstract

This brief studies the discounted stochastic linear quadratic regulator (LQR) problem for systems suffering from additive noise of unknown mean. A completely model-free (MF) value iteration (VI) algorithm is developed to learn the optimal control policy using off-line system trajectories. The generated control policies are proven to converge to a small neighborhood of the optimal ones with high probability. In addition, an MF algorithm is proposed to learn a feasible discount factor. The proposed MF algorithms are illustrated through several examples.