Object pose estimation is a critical task in robotics for precise object manipulation. However, current techniques heavily rely on a reference 3D object, limiting their generalizability and making it expensive to expand to new object categories. Direct pose predictions also provide limited information for robotic grasping without referencing the 3D model. Keypoint-based methods offer intrinsic descriptiveness without relying on an exact 3D model, but they may lack consistency and accuracy. To address these challenges, this paper proposes ShapeShift, a superquadric-based framework for object pose estimation that predicts the object's pose relative to a primitive shape which is fitted to the object. The proposed framework offers intrinsic descriptiveness and the ability to generalize to arbitrary geometric shapes beyond the training set.
翻译:物体姿态估计是机器人精确操作物体中的关键任务。然而,当前技术严重依赖参考3D物体,限制了其泛化能力,且扩展到新物体类别成本高昂。直接姿态预测在不参考3D模型的情况下,为机器人抓取提供的信息有限。基于关键点的方法无需精确的3D模型即可提供内在描述性,但其一致性和准确性可能不足。为解决这些挑战,本文提出ShapeShift,一种基于超二次曲面的物体姿态估计框架,它预测物体相对于拟合到物体的原始形状的姿态。该框架具有内在描述性,并能泛化到训练集之外的任意几何形状。