Ensemble Encoder-Enabled Proactive Human Assembly Intention Recognition With Multimodal and Flexible Scale Data.
Journal:
IEEE transactions on cybernetics
Published Date:
May 1, 2026
Abstract
Human-robot collaboration (HRC) assembly necessitates precise mutual cognition to guarantee safe and efficient execution. In this context, human assembly intention recognition (HAIR) serves as a critical approach to achieving this mutual understanding. However, most current HAIR approaches struggle to extract sufficient spatiotemporal information from limited industrial data, particularly under complex conditions like varying scales and visual occlusions. Thereby, this article proposes an ensemble encoder approach to extract and fuse spatial and temporal features from visual and skeleton streams of the HRC assembly process, thus significantly improving HAIR accuracy and efficiency. First, an RGB feature extraction encoder is designed to model spatiotemporal dependencies of the assembly process with different scales of features from flexible input RGB encoders (RGBEs). Distinctively, a cross-attention module is utilized to fuse information from different-scale RGBEs, ensuring comprehensive assembly action representation with different granularities. Second, to address the occlusion challenge, a mask-aware skeleton feature extraction encoder is devised. By utilizing frame and joint masking strategies, it robustly models the relationship between operator pose evolution and assembly actions, maintaining high performance even under occlusion. Third, a global feature fusion encoder integrates and aligns features from RGB and skeleton feature extraction encoders. Experimental results demonstrate the state-of-the-art performance of the proposed approach, which achieves the highest accuracy of 99.12%, 99.23%, and 84.59% on MCV-Intention, HA4M, and HA-VID datasets, respectively. Six ablation studies demonstrate the performance effects of fusion positions, the number of depth channels, cross-attention fusion module, occlusions, illuminations, and computational efficiency.
Authors
Keywords
No keywords available for this article.