The analysis and use of egocentric videos for robotic tasks is made challenging by occlusion due to the hand and the visual mismatch between the human hand and a robot end-effector. In this sense, the human hand presents a nuisance. However, often hands also provide a valuable signal, e.g. the hand pose may suggest what kind of object is being held. In this work, we propose to extract a factored representation of the scene that separates the agent (human hand) and the environment. This alleviates both occlusion and mismatch while preserving the signal, thereby easing the design of models for downstream robotics tasks. At the heart of this factorization is our proposed Video Inpainting via Diffusion Model (VIDM) that leverages both a prior on real-world images (through a large-scale pre-trained diffusion model) and the appearance of the object in earlier frames of the video (through attention). Our experiments demonstrate the effectiveness of VIDM at improving inpainting quality on egocentric videos and the power of our factored representation for numerous tasks: object detection, 3D reconstruction of manipulated objects, and learning of reward functions, policies, and affordances from videos.
翻译:第一人称视频在机器人任务中的分析和应用面临两大挑战:手部造成的遮挡,以及人类手部与机器人末端执行器之间的视觉差异。从这一角度看,人类手部构成了干扰。然而,手部往往也传递着宝贵信息,例如手部姿态可能暗示所持物体的类型。本文提出一种场景的分解式表征方法,将智能体(人类手部)与环境分离开来。这种方法在保留信号的同时缓解了遮挡与视觉差异问题,从而简化了下游机器人任务模型的设计。该分解框架的核心是我们提出的基于扩散模型的视频修复方法(VIDM),该方法既利用大规模预训练扩散模型对真实世界图像的先验知识,又通过注意力机制利用视频早期帧中物体的外观信息。实验证明,VIDM在提升第一人称视频修复质量方面效果显著,而我们的分解式表征在多项任务中展现出强大能力:目标检测、操作物体的三维重建,以及从视频中学习奖励函数、策略与可供性。