Aligning Foundation Model Priors and Diffusion-Based Hand Interactions for Occlusion-Resistant Two-Hand Reconstruction
Journal:
arXiv
Published Date:
Mar 22, 2025
Abstract
Two-hand reconstruction from monocular images faces persistent challenges due
to complex and dynamic hand postures and occlusions, causing significant
difficulty in achieving plausible interaction alignment. Existing approaches
struggle with such alignment issues, often resulting in misalignment and
penetration artifacts. To tackle this, we propose a novel framework that
attempts to precisely align hand poses and interactions by synergistically
integrating foundation model-driven 2D priors with diffusion-based interaction
refinement for occlusion-resistant two-hand reconstruction. First, we introduce
a Fusion Alignment Encoder that learns to align fused multimodal priors
keypoints, segmentation maps, and depth cues from foundation models during
training. This provides robust structured guidance, further enabling efficient
inference without foundation models at test time while maintaining high
reconstruction accuracy. Second, we employ a two-hand diffusion model
explicitly trained to transform interpenetrated poses into plausible,
non-penetrated interactions, leveraging gradient-guided denoising to correct
artifacts and ensure realistic spatial relations. Extensive evaluations
demonstrate that our method achieves state-of-the-art performance on
InterHand2.6M, FreiHAND, and HIC datasets, significantly advancing occlusion
handling and interaction robustness.