We present Universal Manipulation Interface (UMI) -- a data collection and policy learning framework that allows direct skill transfer from in-the-wild human demonstrations to deployable robot policies. UMI employs hand-held grippers coupled with careful interface design to enable portable, low-cost, and information-rich data collection for challenging bimanual and dynamic manipulation demonstrations. To facilitate deployable policy learning, UMI incorporates a carefully designed policy interface with inference-time latency matching and a relative-trajectory action representation. The resulting learned policies are hardware-agnostic and deployable across multiple robot platforms. Equipped with these features, UMI framework unlocks new robot manipulation capabilities, allowing zero-shot generalizable dynamic, bimanual, precise, and long-horizon behaviors, by only changing the training data for each task. We demonstrate UMI's versatility and efficacy with comprehensive real-world experiments, where policies learned via UMI zero-shot generalize to novel environments and objects when trained on diverse human demonstrations. UMI's hardware and software system is open-sourced at https://umi-gripper.github.io.
翻译:我们提出通用操控接口(UMI)——一种数据采集与策略学习框架,可直接将野外人类演示的技能迁移至可部署的机器人策略。UMI采用手持夹爪配合精心设计的接口方案,实现便携、低成本且信息丰富的挑战性双手及动态操控演示数据采集。为促进可部署策略学习,UMI集成了经过推理时延匹配与相对轨迹动作表示精心设计的策略接口。由此习得的策略具有硬件无关性,可跨多种机器人平台部署。凭借这些特性,UMI框架解锁了新的机器人操控能力,仅需为每项任务更换训练数据,即可实现零样本泛化的动态、双手、精准及长时程行为。通过全面的真实世界实验,我们验证了UMI的通用性与有效性:基于多样化人类演示训练的UMI学习策略,在零样本条件下即可泛化至全新环境与物体。UMI的硬件与软件系统已在https://umi-gripper.github.io开源。