Policy-Adjustable Q-Learning for Data-Driven Nonlinear Optimal Tracking Control.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41838502.
- Also identified by DOI 10.1109/TNNLS.2026.3672136.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
This article investigates a novel policy-adjustable Q-learning (PA-QL) algorithm aimed at addressing the optimal tracking control (OTC) problem for nonlinear discrete-time (DT) systems with enhanced adaptability and flexibility. A novel iteration scheme is developed that integrates the control weights into the augmented neural network (NN) input, thereby reformulating the learning process to explicitly characterize the optimal policy as a function of the adjustable weights. Consequently, the learned control policy is not constrained by predetermined weights, allowing for dynamic adjustment after offline training is completed. Moreover, such adjustments can be performed online seamlessly, offering substantially greater flexibility in adapting to changes in operating conditions or control objectives. Finally, the effectiveness of the proposed algorithm is established through rigorous theoretical analysis and further validated by simulation studies.