Perception is an essential part of robotic manipulation in a semi-structured environment. Traditional approaches produce a narrow task-specific prediction (e.g., object's 6D pose), that cannot be adapted to other tasks and is ill-suited for deformable objects. In this paper, we propose using canonical mapping as a near-universal and flexible object descriptor. We demonstrate that common object representations can be derived from a single pre-trained canonical mapping model, which in turn can be generated with minimal manual effort using an automated data generation and training pipeline. We perform a multi-stage experiment using two robot arms that demonstrate the robustness of the perception approach and the ways it can inform the manipulation strategy, thus serving as a powerful foundation for general-purpose robotic manipulation.
翻译:感知是半结构化环境中机器人操作的关键组成部分。传统方法通常产生狭窄的、面向特定任务的预测(例如物体的六自由度位姿),这些预测无法适应其他任务,且不适用于可变形物体。本文提出将规范映射作为近乎通用的、灵活的物体描述符。我们证明,常见的物体表征可从单个预训练的规范映射模型推导得出,而该模型可通过自动化数据生成和训练流程以最小的人工投入生成。我们采用两台机械臂进行了多阶段实验,验证了该感知方法的鲁棒性及其对操作策略的指导方式,从而奠定通用机器人操作的强大基础。