In this work, we present 3DCoMPaT$^{++}$, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the part-instance level, alongside matching RGB point clouds, 3D textured meshes, depth maps, and segmentation masks. 3DCoMPaT$^{++}$ covers 41 shape categories, 275 fine-grained part categories, and 293 fine-grained material classes that can be compositionally applied to parts of 3D objects. We render a subset of one million stylized shapes from four equally spaced views as well as four randomized views, leading to a total of 160 million renderings. Parts are segmented at the instance level, with coarse-grained and fine-grained semantic levels. We introduce a new task, called Grounded CoMPaT Recognition (GCR), to collectively recognize and ground compositions of materials on parts of 3D objects. Additionally, we report the outcomes of a data challenge organized at CVPR2023, showcasing the winning method's utilization of a modified PointNet$^{++}$ model trained on 6D inputs, and exploring alternative techniques for GCR enhancement. We hope our work will help ease future research on compositional 3D Vision.
翻译:本文提出3DCoMPaT$^{++}$,一个包含1.6亿渲染视图的多模态2D/3D数据集,覆盖超过1000万个精细部件级标注的风格化三维模型,同时提供配对的RGB点云、带纹理的三维网格、深度图及分割掩膜。该数据集涵盖41种形状类别、275种细粒度部件类别和293种可组合应用于三维物体部件的细粒度材质类别。我们从四个等距视角及四个随机视角对100万个风格化形状子集进行渲染,共计生成1.6亿张渲染图。部件分割达到实例级别,并包含粗粒度和细粒度两个语义层级。我们提出了一项名为"接地组合识别(GCR)"的新任务,旨在共同识别并定位三维物体部件上的材质组合。此外,我们汇报了CVPR2023组织的数据挑战赛成果,展示了优胜方法如何利用改进的PointNet$^{++}$模型处理6D输入,并探索了增强GCR性能的替代技术。希望本研究能促进未来三维视觉组合识别领域的研究进展。