Efficient multi-agent policy adaptation with Bayesian policy reuse and view-invariant contrastive awareness.
Journal:
Neural networks : the official journal of the International Neural Network Society
Published Date:
Feb 21, 2026
Abstract
Bayesian Policy Reuse (BPR), a framework that selects from a library of pre-trained policies via Bayesian belief updates, enables efficient response to non-stationary opponents who suddenly switch strategies. However, most previous work is restricted to single controllable agent settings where partial observability is also absent. In this paper, we propose BPR-VCA, which integrates BPR with a multi-view contrastive learning module to achieve efficient policy adaptation in partially observable multi-agent environments, where each controllable agent only accesses local observation trajectories (i.e., its own observation-action history). To address limited observability, unlike prior BPR methods that rely on global opponent trajectories or episodic rewards, we use local trajectories as observation signals. To tackle the partial observability, in contrast to existing BPR-based algorithms that use opponent trajectories or episodic rewards as dependent information, we adopt local trajectories as observation signals. Moreover, a View-invariant Contrastive Awareness (VCA) module is designed to facilitates cognitive consensus among the controllable agents about global dynamics changes. Specifically, it integrates global and local views of the same task, while ensuring that views of different tasks are distinctly separated. Leveraging the acquired contextual features, we establish local observation models for decentralized online beliefs updating. During online decentralized execution, the controllable agents update their beliefs respectively and finally can reuse the most appropriate response joint policies. Experiments on four competitive scenarios show that BPR-VCA achieves higher episodic and accumulated rewards, faster and more accurate opponent recognition, and higher win rates compared with state-of-the-art baselines.
Authors
Keywords
No keywords available for this article.