Decomposition and transfer of individual Q-values for decision-making of multi-agent reinforcement learning with communication.

Journal: Neural networks : the official journal of the International Neural Network Society
Published Date:

Abstract

In partially observable multi-agent tasks, communication modules need to be trained to supplement information for decision-making, and decision-making modules need to be trained to analyze received information. This leads to their coupling during the process of multi-agent reinforcement learning. To alleviate the difficulty from the coupling, two transfer-based frameworks are proposed. The Simple Transfer Framework converts a difficult reinforcement learning into an easy one and a supervised learning. The Decomposition and Transfer Framework further decomposes this supervised learning into two simpler ones, based on the decomposition of communication-based individual Q-values. They are decomposed into local individual Q-values and their deviations. Based on this framework, the mutual information neural estimation is introduced to train message generators independently. The proposed frameworks are verified by training two communication-based architectures. They outperform direct reinforcement learning algorithms across various team sizes tested in the Drone Fight. The independent training method effectively facilitates the transfer of the deviations.

Authors

Keywords

No keywords available for this article.