Supervised learning models for precise tracking of hand-object interactions (HOI) in 3D require large amounts of annotated data for training. Moreover, it is not intuitive for non-experts to label 3D ground truth (e.g. 6DoF object pose) on 2D images. To address these issues, we present "blender-hoisynth", an interactive synthetic data generator based on the Blender software. Blender-hoisynth can scalably generate and automatically annotate visual HOI training data. Other competing approaches usually generate synthetic HOI data compeletely without human input. While this may be beneficial in some scenarios, HOI applications inherently necessitate direct control over the HOIs as an expression of human intent. With blender-hoisynth, it is possible for users to interact with objects via virtual hands using standard Virtual Reality hardware. The synthetically generated data are characterized by a high degree of photorealism and contain visually plausible and physically realistic videos of hands grasping objects and moving them around in 3D. To demonstrate the efficacy of our data generation, we replace large parts of the training data in the well-known DexYCB dataset with hoisynth data and train a state-of-the-art HOI reconstruction model with it. We show that there is no significant degradation in the model performance despite the data replacement.
翻译:用于精确追踪三维空间中手-物体交互的有监督学习模型需要大量标注数据进行训练。此外,非专家人员在二维图像上标注三维真实值(如六自由度物体位姿)并非直观易行。为解决这些问题,我们提出基于Blender软件的交互式合成数据生成器"blender-hoisynth"。该工具可规模化生成并自动标注视觉手-物体交互训练数据。其他同类方法通常完全脱离人工输入生成合成手-物体交互数据,虽然这在某些场景下具有优势,但手-物体交互应用本质上需要将交互过程作为人类意图的表达进行直接控制。借助blender-hoisynth,用户可通过标准虚拟现实硬件使用虚拟手与物体进行交互。所生成的合成数据具有高度照片级真实感,包含视觉合理且物理真实的手部抓取物体并在三维空间中移动的视频片段。为验证数据生成效能,我们将知名DexYCB数据集中的大部分训练数据替换为hoisynth数据,并训练了最先进的手-物体交互重建模型。结果表明,尽管进行了数据替换,模型性能并未出现显著下降。