Understanding how humans interact with the world necessitates accurate 3D hand pose estimation, a task complicated by the hand's high degree of articulation, frequent occlusions, self-occlusions, and rapid motions. While most existing methods rely on single-image inputs, videos have useful cues to address aforementioned issues. However, existing video-based 3D hand datasets are insufficient for training feedforward models to generalize to in-the-wild scenarios. On the other hand, we have access to large human motion capture datasets which also include hand motions, e.g. AMASS. Therefore, we develop a generative motion prior specific for hands, trained on the AMASS dataset which features diverse and high-quality hand motions. This motion prior is then employed for video-based 3D hand motion estimation following a latent optimization approach. Our integration of a robust motion prior significantly enhances performance, especially in occluded scenarios. It produces stable, temporally consistent results that surpass conventional single-frame methods. We demonstrate our method's efficacy via qualitative and quantitative evaluations on the HO3D and DexYCB datasets, with special emphasis on an occlusion-focused subset of HO3D. Code is available at https://hmp.is.tue.mpg.de
翻译:理解人类与世界的交互需要精确的3D手部姿态估计,然而手部的高自由度运动、频繁的遮挡、自遮挡以及快速动作使得该任务极具挑战性。现有方法大多依赖单张图像输入,而视频数据包含解决上述问题的有效线索。然而,现有的基于视频的3D手部数据集不足以训练前馈模型以泛化到野外场景。另一方面,我们已获取包含手部运动的大规模人体运动捕捉数据集(如AMASS)。为此,我们基于包含多样化高质量手部运动的AMASS数据集,提出了一种专门针对手部的生成式运动先验。该运动先验通过潜在优化方法被应用于基于视频的3D手部运动估计。我们融合的鲁棒运动先验显著提升了性能,尤其在遮挡场景中表现突出,生成了比传统单帧方法更稳定、更时序一致的结果。通过在HO3D和DexYCB数据集上的定性与定量评估,并特别关注HO3D的遮挡聚焦子集,我们验证了该方法的有效性。代码已开源:https://hmp.is.tue.mpg.de