Humans possess the cognitive ability to comprehend scenes in a compositional manner. To empower AI systems with similar abilities, object-centric representation learning aims to acquire representations of individual objects from visual scenes without any supervision. Although recent advancements in object-centric representation learning have achieved remarkable progress on complex synthesis datasets, there is a huge challenge for application in complex real-world scenes. One of the essential reasons is the scarcity of real-world datasets specifically tailored to object-centric representation learning methods. To solve this problem, we propose a versatile real-world dataset of tabletop scenes for object-centric learning called OCTScenes, which is meticulously designed to serve as a benchmark for comparing, evaluating and analyzing object-centric representation learning methods. OCTScenes contains 5000 tabletop scenes with a total of 15 everyday objects. Each scene is captured in 60 frames covering a 360-degree perspective. Consequently, OCTScenes is a versatile benchmark dataset that can simultaneously satisfy the evaluation of object-centric representation learning methods across static scenes, dynamic scenes, and multi-view scenes tasks. Extensive experiments of object-centric representation learning methods for static, dynamic and multi-view scenes are conducted on OCTScenes. The results demonstrate the shortcomings of state-of-the-art methods for learning meaningful representations from real-world data, despite their impressive performance on complex synthesis datasets. Furthermore, OCTScenes can serves as a catalyst for advancing existing state-of-the-art methods, inspiring them to adapt to real-world scenes. Dataset and code are available at https://huggingface.co/datasets/Yinxuan/OCTScenes.
翻译:人类具备以组合方式理解场景的认知能力。为使人工智能系统获得类似能力,目标中心表征学习旨在无需任何监督的情况下从视觉场景中获取单个目标的表征。尽管近年来目标中心表征学习在复杂合成数据集上取得了显著进展,但在复杂真实场景中的应用仍然面临巨大挑战。其中一个关键原因是专门针对目标中心表征学习方法设计的真实世界数据集十分匮乏。为解决该问题,我们提出了一个名为OCTScenes的多用途桌面场景真实世界数据集,该数据集经过精心设计,可作为比较、评估与分析目标中心表征学习方法的基准。OCTScenes包含5000个桌面场景,涵盖15种日常物品,每个场景通过60帧图像覆盖360度视角。因此,OCTScenes是一个多用途基准数据集,能够同时满足静态场景、动态场景及多视角场景任务中目标中心表征学习方法的评估需求。我们在OCTScenes上对面向静态、动态及多视角场景的目标中心表征学习方法进行了大量实验。结果表明,尽管现有最先进方法在复杂合成数据集上表现优异,但在真实世界数据中学习有意义的表征仍存在不足。此外,OCTScenes可推动现有最先进方法的改进,促使其适应真实场景。数据集与代码已发布在https://huggingface.co/datasets/Yinxuan/OCTScenes。