Multi-agent contrastive exploration via value decomposition discrepancy.
Journal:
Neural networks : the official journal of the International Neural Network Society
Published Date:
Apr 6, 2026
Abstract
Value decomposition, as a factorization approach in multi-agent reinforcement learning (MARL), has been influential in the development of many effective value-based algorithms. Existing studies on value decomposition that focus on representation capability often suffer from low sample efficiency, while other factorization approaches may also have limited representation capability, which may still prevent collaborators from discovering the optimal joint action. To enhance agents' exploration with an unlimited joint state-value function, we propose a Multi-Agent Contrastive Exploration method (MACE) leveraging value decomposition discrepancies and contrastive principles. MACE determines update weights based on the discrepancy between the different value decomposition estimates to set update weights and introduces this difference as an intrinsic target in the update process. Additionally, MACE designs an exploration preference network inspired by this difference, explicitly adjusting the exploration preferences of agents during interactions. Through the experiments of various types and difficulties on Matrix Games and Starcraft Multi-Agent Challenge, we show that MACE not only significantly outperforms the baselines in learning speed and final performance, but also effectively maintains the higher expressivity, representing an innovative solution that integrates the advantages of existing algorithms.
Authors
Keywords
No keywords available for this article.