We tackle the task of reconstructing hand-object interactions from short video clips. Given an input video, our approach casts 3D inference as a per-video optimization and recovers a neural 3D representation of the object shape, as well as the time-varying motion and hand articulation. While the input video naturally provides some multi-view cues to guide 3D inference, these are insufficient on their own due to occlusions and limited viewpoint variations. To obtain accurate 3D, we augment the multi-view signals with generic data-driven priors to guide reconstruction. Specifically, we learn a diffusion network to model the conditional distribution of (geometric) renderings of objects conditioned on hand configuration and category label, and leverage it as a prior to guide the novel-view renderings of the reconstructed scene. We empirically evaluate our approach on egocentric videos across 6 object categories, and observe significant improvements over prior single-view and multi-view methods. Finally, we demonstrate our system's ability to reconstruct arbitrary clips from YouTube, showing both 1st and 3rd person interactions.
翻译:我们致力于从短视频片段中重建手物交互任务。给定输入视频,本方法将三维推理转化为逐视频优化过程,恢复物体形状的神经三维表示,以及时变运动与手部关节姿态。尽管输入视频自然提供了部分多视角线索以指导三维推理,但由于遮挡和视角变化有限,这些线索本身并不充分。为获取精确的三维信息,我们通过通用数据驱动先验增强多视角信号以引导重建。具体而言,我们学习一个扩散网络来建模物体(几何)渲染在给定手部配置与类别标签条件下的条件分布,并将其作为先验指导重建场景的新视角渲染。我们在涵盖6类物体类别的第一人称视频上进行了实证评估,观察到此方法较先前的单视角与多视角方法有显著提升。最后,我们展示了系统从YouTube重建任意片段的能力,包括第一人称和第三人称交互场景。