Multi-agent contrastive exploration via value decomposition discrepancy.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41980548.
- Also identified by DOI 10.1016/j.neunet.2026.108929.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Value decomposition, as a factorization approach in multi-agent reinforcement learning (MARL), has been influential in the development of many effective value-based algorithms. Existing studies on value decomposition that focus on representation capability often suffer from low sample efficiency, while other factorization approaches may also have limited representation capability, which may still prevent collaborators from discovering the optimal joint action. To enhance agents' exploration with an unlimited joint state-value function, we propose a Multi-Agent Contrastive Exploration method (MACE) leveraging value decomposition discrepancies and contrastive principles. MACE determines update weights based on the discrepancy between the different value decomposition estimates to set update weights and introduces this difference as an intrinsic target in the update process. Additionally, MACE designs an exploration preference network inspired by this difference, explicitly adjusting the exploration preferences of agents during interactions. Through the experiments of various types and difficulties on Matrix Games and Starcraft Multi-Agent Challenge, we show that MACE not only significantly outperforms the baselines in learning speed and final performance, but also effectively maintains the higher expressivity, representing an innovative solution that integrates the advantages of existing algorithms.