Online Reinforcement Learning Control Designs With Acceleration Mechanism for Unknown Multiagent Systems Through Value Iteration.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40366832.
- Also identified by DOI 10.1109/TNNLS.2025.3563155.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In this article, an online reinforcement learning (RL) control method through value iteration (VI) is developed to solve the optimal cooperative control problem for the unknown linear discrete-time multiagent systems (MASs). On the one hand, an online learning scheme with evolving policies is proposed in order to guarantee the stability of the MASs under immature policies generated by VI. Inspired by the event-triggered mechanism, the stability criterion is designed as a trigger to filter the admissible control policies, which eliminates the need to establish a monotonic value function sequence. On the other hand, an acceleration mechanism for the MASs is presented such that the convergence rate of VI can be accelerated. The relationship between the selection of the relaxation factor and the accelerated convergence process is elaborated. Simple backpropagation (BP) neural networks (NNs) are applied for the implementation. Two classical examples are introduced and simulation results are provided in order to substantiate the validity of the designed method.