Online Adaptive Optimal Control Algorithm Based on Weighted Policy Iteration.

Tan, Wanlin; Luo, Rui; Peng, Zhinan; Ling, Qiang · IEEE Trans Neural Netw Learn Syst · 2025

basic_science · Level V

Where this comes from

Abstract

In this article, we propose a novel online learning algorithm based on weighted policy iteration (WPI) for addressing optimal control problems of nonlinear systems. WPI is proposed to deal with the influence of the neural network (NN) approximation error on the admissibility of the improved control policy. It is shown that the new iterative method can converge to the optimal solution uniformly. Utilizing NN approximation and experience replaying techniques, a WPI-based online learning algorithm is proposed. The new online algorithm distinguishes from previously known ones in that there can be fewer neurons in the hidden layer, giving rise to significant computational improvement. The assumption that the number of neurons needs to be sufficiently large can be dropped. Besides, instead of a standard persistently excited (PE) condition, only a relaxed PE condition is needed, which is also easier to check. Finally, numerical experiments are conducted to verify the effectiveness of our method.