We propose OmniNOCS, a large-scale monocular dataset with 3D Normalized Object Coordinate Space (NOCS) maps, object masks, and 3D bounding box annotations for indoor and outdoor scenes. OmniNOCS has 20 times more object classes and 200 times more instances than existing NOCS datasets (NOCS-Real275, Wild6D). We use OmniNOCS to train a novel, transformer-based monocular NOCS prediction model (NOCSformer) that can predict accurate NOCS, instance masks and poses from 2D object detections across diverse classes. It is the first NOCS model that can generalize to a broad range of classes when prompted with 2D boxes. We evaluate our model on the task of 3D oriented bounding box prediction, where it achieves comparable results to state-of-the-art 3D detection methods such as Cube R-CNN. Unlike other 3D detection methods, our model also provides detailed and accurate 3D object shape and segmentation. We propose a novel benchmark for the task of NOCS prediction based on OmniNOCS, which we hope will serve as a useful baseline for future work in this area. Our dataset and code will be at the project website: https://omninocs.github.io.
翻译:我们提出了OmniNOCS,这是一个大规模的单目数据集,包含用于室内外场景的3D归一化物体坐标空间(NOCS)图、物体掩码和3D边界框标注。OmniNOCS所包含的物体类别是现有NOCS数据集(NOCS-Real275、Wild6D)的20倍,实例数量是其200倍。我们利用OmniNOCS训练了一个新颖的、基于Transformer的单目NOCS预测模型(NOCSformer),该模型能够从跨多样类别的2D物体检测中预测精确的NOCS、实例掩码和姿态。这是首个在给定2D边界框提示时,能够泛化到广泛类别的NOCS模型。我们在3D定向边界框预测任务上评估了我们的模型,其性能与Cube R-CNN等最先进的3D检测方法相当。与其他3D检测方法不同,我们的模型还能提供详细且精确的3D物体形状和分割结果。我们基于OmniNOCS为NOCS预测任务提出了一个新的基准测试,希望它能作为该领域未来工作的有用基线。我们的数据集和代码将发布在项目网站:https://omninocs.github.io。