After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning
Chaoyuan Hao, Wentao Yue, Tianyou Lai, Hongji Li, Jiayi Zhou, Qingyu Mao, Qilei Li
Abstract
Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.
Create a lesson
Related papers
Managing Context and Communication in Distributed Agentic UAV Swarms
Andrea Iannoli, Ivan Zyrianoff, Angelo Trotta et al.
LLM-Driven Multi-Agent Control for Skill-Based Smart Manufacturing
Kay Köhle, Darko Anicic, Thomas A. Runkler et al.
Can AI Scientists Coordinate at Runtime?
Zijian Liu, Yangzhixin Luo, Junyu Lu et al.
Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
Yunbei Zhang, Saiyue Lyu, Janet Wang et al.
Solving Multi-Agent Sokoban via LaCAM
Keisuke Okumura
Consensus and Factual Dynamics in Large Populations of Interacting Language Models
Emanuele Ricco, Elia Onofri, Vincenzo Sammartino et al.