We present CARTO, a novel approach for reconstructing multiple articulated objects from a single stereo RGB observation. We use implicit object-centric representations and learn a single geometry and articulation decoder for multiple object categories. Despite training on multiple categories, our decoder achieves a comparable reconstruction accuracy to methods that train bespoke decoders separately for each category. Combined with our stereo image encoder we infer the 3D shape, 6D pose, size, joint type, and the joint state of multiple unknown objects in a single forward pass. Our method achieves a 20.4% absolute improvement in mAP 3D IOU50 for novel instances when compared to a two-stage pipeline. Inference time is fast and can run on a NVIDIA TITAN XP GPU at 1 HZ for eight or less objects present. While only trained on simulated data, CARTO transfers to real-world object instances. Code and evaluation data is available at: http://carto.cs.uni-freiburg.de
翻译:我们提出CARTO,一种从单次立体RGB观测中重建多个铰接物体的新方法。我们采用隐式物体中心表示,并为多个物体类别学习单一几何与关节解码器。尽管在多类别数据上训练,我们的解码器在重建精度上仍能达到与为每个类别单独训练定制解码器的方法相当的水平。结合我们的立体图像编码器,可在单次前向传播中推断多个未知物体的3D形状、6D位姿、尺寸、关节类型及关节状态。与两阶段流水线相比,我们的方法在新实例上的mAP 3D IOU50指标实现了20.4%的绝对提升。推理速度快,在NVIDIA TITAN XP GPU上对于八个或更少物体可达1Hz处理频率。尽管仅在模拟数据上训练,CARTO可迁移至真实世界物体实例。代码与评估数据见:http://carto.cs.uni-freiburg.de