Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, yet such data is scarce. Real-world capture is costly and limited to controlled settings, while existing synthetic datasets rely on simple geometric heuristics, ignoring rich scene context. In contrast, 2D foundation models trained at internet scale have acquired commonsense knowledge of human-environment interactions. To transfer this knowledge to 3D, we introduce InHabit, an automatic and scalable data generator for populating 3D scenes with interacting humans. InHabit follows a render-generate-lift principle: given a rendered 3D scene, a vision-language model proposes contextually meaningful actions, an image-editing model inserts a human, and an optimization procedure lifts the edited result into physically plausible SMPL-X bodies aligned with the scene geometry. Applied to Habitat-Matterport3D, InHabit produces InHabitants, the first large-scale photorealistic 3D human-scene interaction dataset, with 78K samples across $\sim$800 building-scale scenes with complete 3D geometry, SMPL-X bodies, and images. Augmenting standard training data with InHabitants improves RGB-based 3D human-scene reconstruction and contact estimation, and in a perceptual user study our data is preferred in 78% of cases over prior art.


翻译:摘要:为了让具身智能体像人类一样理解三维场景,需要大量关于人与多样化环境进行有意义交互的数据,然而此类数据十分稀缺。真实世界的数据采集成本高昂且局限于受控环境,而现存的合成数据集则依赖简单的几何启发式规则,忽略了丰富的场景上下文。相比之下,在互联网规模上训练的二维基础模型已积累了关于人-环境交互的常识知识。为了将这些知识迁移至三维领域,我们提出了InHabit——一个用于在三维场景中填充交互人体的自动化、可扩展数据生成器。InHabit遵循“渲染-生成-提升”原则:给定一个渲染后的三维场景,首先由视觉-语言模型提出具有上下文意义的动作,再由图像编辑模型插入人体,最后通过优化过程将编辑结果提升为物理上合理的SMPL-X人体模型,使其与场景几何对齐。将InHabit应用于Habitat-Matterport3D数据集后,生成了InHabitants——首个大规模、逼真的三维人-场景交互数据集,包含约800个建筑级场景中的78K个样本,并提供了完整的三维几何、SMPL-X人体模型及图像信息。将InHabitants作为标准训练数据的补充,能够提升基于RGB的三维人体-场景重建与接触估计性能;在感知用户研究中,我们的数据在78%的情况下优于先前方法。

0
下载
关闭预览

相关内容

3D是英文“Three Dimensions”的简称,中文是指三维、三个维度、三个坐标,即有长、有宽、有高,换句话说,就是立体的,是相对于只有长和宽的平面(2D)而言。
面向具身智能与机器人仿真的三维生成:综述
专知会员服务
18+阅读 · 4月30日
具身智能中的心理世界建模:深度综述
专知会员服务
39+阅读 · 1月10日
三维与四维世界建模综述
专知会员服务
31+阅读 · 2025年9月12日
三维场景生成:综述
专知会员服务
21+阅读 · 2025年5月9日
三维物体与场景生成的最新进展:综述
专知会员服务
19+阅读 · 2025年4月17日
如何构建行业知识图谱(以医疗行业为例)
SkeletonNet:完整的人体三维位姿重建方法
计算机视觉life
21+阅读 · 2019年1月21日
【知识图谱】基于知识图谱的用户画像技术
产业智能官
103+阅读 · 2019年1月9日
深度学习时代的图模型,清华发文综述图网络
GAN生成式对抗网络
13+阅读 · 2018年12月23日
报名 | 让机器读懂你的意图——人体姿态估计入门
人工智能头条
10+阅读 · 2017年9月19日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
7+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
0+阅读 · 22分钟前
《履带式无人地面战车技术发展现状》
专知会员服务
1+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关VIP内容
面向具身智能与机器人仿真的三维生成:综述
专知会员服务
18+阅读 · 4月30日
具身智能中的心理世界建模:深度综述
专知会员服务
39+阅读 · 1月10日
三维与四维世界建模综述
专知会员服务
31+阅读 · 2025年9月12日
三维场景生成:综述
专知会员服务
21+阅读 · 2025年5月9日
三维物体与场景生成的最新进展:综述
专知会员服务
19+阅读 · 2025年4月17日
相关资讯
相关基金
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
7+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员