Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, yet such data is scarce. Real-world capture is costly and limited to controlled settings, while existing synthetic datasets rely on simple geometric heuristics, ignoring rich scene context. In contrast, 2D foundation models trained at internet scale have acquired commonsense knowledge of human-environment interactions. To transfer this knowledge to 3D, we introduce InHabit, an automatic and scalable data generator for populating 3D scenes with interacting humans. InHabit follows a render-generate-lift principle: given a rendered 3D scene, a vision-language model proposes contextually meaningful actions, an image-editing model inserts a human, and an optimization procedure lifts the edited result into physically plausible SMPL-X bodies aligned with the scene geometry. Applied to Habitat-Matterport3D, InHabit produces InHabitants, the first large-scale photorealistic 3D human-scene interaction dataset, with 78K samples across $\sim$800 building-scale scenes with complete 3D geometry, SMPL-X bodies, and images. Augmenting standard training data with InHabitants improves RGB-based 3D human-scene reconstruction and contact estimation, and in a perceptual user study our data is preferred in 78% of cases over prior art.
翻译:摘要:为了让具身智能体像人类一样理解三维场景,需要大量关于人与多样化环境进行有意义交互的数据,然而此类数据十分稀缺。真实世界的数据采集成本高昂且局限于受控环境,而现存的合成数据集则依赖简单的几何启发式规则,忽略了丰富的场景上下文。相比之下,在互联网规模上训练的二维基础模型已积累了关于人-环境交互的常识知识。为了将这些知识迁移至三维领域,我们提出了InHabit——一个用于在三维场景中填充交互人体的自动化、可扩展数据生成器。InHabit遵循“渲染-生成-提升”原则:给定一个渲染后的三维场景,首先由视觉-语言模型提出具有上下文意义的动作,再由图像编辑模型插入人体,最后通过优化过程将编辑结果提升为物理上合理的SMPL-X人体模型,使其与场景几何对齐。将InHabit应用于Habitat-Matterport3D数据集后,生成了InHabitants——首个大规模、逼真的三维人-场景交互数据集,包含约800个建筑级场景中的78K个样本,并提供了完整的三维几何、SMPL-X人体模型及图像信息。将InHabitants作为标准训练数据的补充,能够提升基于RGB的三维人体-场景重建与接触估计性能;在感知用户研究中,我们的数据在78%的情况下优于先前方法。