We contribute the Habitat Synthetic Scene Dataset, a dataset of 211 high-quality 3D scenes, and use it to test navigation agent generalization to realistic 3D environments. Our dataset represents real interiors and contains a diverse set of 18,656 models of real-world objects. We investigate the impact of synthetic 3D scene dataset scale and realism on the task of training embodied agents to find and navigate to objects (ObjectGoal navigation). By comparing to synthetic 3D scene datasets from prior work, we find that scale helps in generalization, but the benefits quickly saturate, making visual fidelity and correlation to real-world scenes more important. Our experiments show that agents trained on our smaller-scale dataset can match or outperform agents trained on much larger datasets. Surprisingly, we observe that agents trained on just 122 scenes from our dataset outperform agents trained on 10,000 scenes from the ProcTHOR-10K dataset in terms of zero-shot generalization in real-world scanned environments.
翻译:我们贡献了栖息地合成场景数据集(Habitat Synthetic Scene Dataset,HSSD),该数据集包含211个高质量三维场景,并用于测试导航智能体在真实三维环境中的泛化能力。本数据集呈现了真实的室内空间,并涵盖18,656个多样化真实世界物体模型。我们系统研究了合成三维场景数据集的规模与真实感对训练具身智能体执行目标查找与导航任务(ObjectGoal navigation)的影响。通过与既往工作中的合成三维场景数据集对比,我们发现场景规模有助于泛化,但其效益会快速饱和,这使得视觉保真度及与现实场景的相关性变得更为关键。实验表明,在较小规模数据集上训练的智能体能够达到甚至超越在更大数据集上训练的同类模型。令人惊讶的是,仅在HSSD数据集中122个场景上训练的智能体,其在真实世界扫描环境中的零样本泛化性能即可超越在ProcTHOR-10K数据集的10,000个场景上训练的智能体。