When to Align: Dynamic Behavior Consistency for Multiagent Systems via Intrinsic Rewards.

Lin, Kunyang; Wang, Yufeng; Chen, Peihao; Zeng, Runhao; Lei, Yinjie; Zhou, Siyuan; Du, Qing; Tan, Mingkui et al. · IEEE Trans Neural Netw Learn Syst · 2025

Where this comes from

Abstract

In multiagent systems, learning optimal behavior policies for individual agents remains a challenging yet crucial task. While recent research has made strides in this area, the issue of when agents should maintain consistent behaviors with one another is still not adequately addressed. This article proposes a novel approach to enable agents to autonomously decide whether their behaviors should align with those of their peers by leveraging intrinsic rewards to optimize their policies. We define behavior consistency as the divergence between the actions taken by two agents given the same observations. To encourage agents to be aware of each other's behaviors, we propose dynamic consistency-based intrinsic reward (DCIR), which guides agents in determining when to synchronize their behaviors. In addition, we introduce a dynamic scaling network (DSN) that provides learnable scaling factors at each time step, enabling agents to dynamically decide the extent of rewarding consistent behavior. Our method is evaluated on environments including Multiagent Particle, Google Research Football, and StarCraft II Micromanagement. Experimental results demonstrate its effectiveness in learning optimal policies.