Hands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotics. Although grasp tracking or object manipulation synthesis can produce coarse hand motion, this kind of motion is inevitably noisy and full of jitter. To address this problem, we propose a data-driven method for coarse motion refinement. First, we design a hand-centric representation to describe the dynamic spatial-temporal relation between hands and objects. Compared to the object-centric representation, our hand-centric representation is straightforward and does not require an ambiguous projection process that converts object-based prediction into hand motion. Second, to capture the dynamic clues of hand-object interaction, we propose a new architecture that models the spatial and temporal structure in a hierarchical manner. Extensive experiments demonstrate that our method outperforms previous methods by a noticeable margin.
翻译:手是人机交互的主要媒介。生成合理的三维手物交互运动对虚拟现实与机器人等应用至关重要。尽管抓取跟踪或物体操作合成可产生粗略的手部运动,但此类运动不可避免地存在噪声和抖动问题。为解决这一难题,我们提出一种数据驱动的粗运动精修方法。首先,我们设计了一种以手部为中心的表示方法,用于描述手与物体之间的动态时空关系。与以物体为中心的表示相比,我们的以手部为中心的表示更直观,无需模棱两可的投影过程将基于物体的预测转换为手部运动。其次,为捕捉手物交互的动态线索,我们提出了一种新架构,以层次化方式建模空间与时间结构。大量实验表明,我们的方法显著优于现有方法。