A unified and efficient training framework for open-ended non-transitive games.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41240599.
- Also identified by DOI 10.1016/j.neunet.2025.108248.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Self-play (SP) and Policy-Space Response Oracles (PSRO) are two fundamental training frameworks for solving Nash equilibrium (NE) in games. SP only performs effectively in transitive games, aiming to progressively derive a consistent winning strategy with an efficient warm start. In open-ended non-transitive games (e.g., Rock-Paper-Scissors), PSRO maintains a policy population and approximates the NE in the meta-game. However, PSRO requires training the policy from scratch in each iteration, making it inefficient in large-scale games. To address these limitations, we propose a Unified and Efficient Play (UEP) framework that combines the strengths of SP and PSRO, enabling the solution of NE in open-ended non-transitive games while benefiting from a warm start in each iteration. In addition, a unified regularized distance metric is proposed to trade off the accuracy of SP and the diversity of PSRO, enhancing the overall training efficiency. We present theoretical evidence that UEP can converge to the NE. Empirically, experiments are conducted in various open-ended games with strong non-transitivity. The results validate the superior performance of UEP in approximating NE and generating robust policies compared to prevailing SP and PSRO variants.
Medical subject headings
- Game Theory