WARGM-PDPG: A dual-phase policy gradient graph mamba neural network for device placement algorithms.
Journal:
Neural networks : the official journal of the International Neural Network Society
Published Date:
Mar 17, 2026
Abstract
As neural networks continue to scale in size and complexity, the computational capabilities of individual devices have become increasingly constrained, necessitating the development of efficient parallel training strategies. This challenge is particularly pronounced in device placement tasks, where optimal resource allocation is critical. This work proposes the WARGM-PDPG strategy, which integrates the Weight Adaptive Residual Graph Mamba Network (WARGMamba) for enhanced graph embedding extraction and the Perturbation-Based Dual-Phase Policy Gradient Algorithm (PDPG) for optimizing device placement strategies. WARGMamba alleviates the over-smoothing issue in traditional Graph Neural Networks (GNNs) and dynamically balances global and local node features via a weight-adaptive residual structure, enriching graph embeddings. For reinforcement learning, PDPG uses a dual-phase update mechanism (internal and external phases) to optimize device placement strategies. The internal phase emphasizes short-term node-level optimization and leverages perturbed data augmentation techniques to distinguish between similar node features. Meanwhile, the external phase evaluates long-term policy performance to ensure convergence toward a globally optimal solution. Experimental results demonstrate that the WARGM-PDPG strategy outperforms existing methods, achieving a 95.89% reduction in device placement computation time and a 10.24% to 28.66% improvement in execution time.
Authors
Keywords
No keywords available for this article.