Many real-world manipulation tasks consist of a series of subtasks that are significantly different from one another. Such long-horizon, complex tasks highlight the potential of dexterous hands, which possess adaptability and versatility, capable of seamlessly transitioning between different modes of functionality without the need for re-grasping or external tools. However, the challenges arise due to the high-dimensional action space of dexterous hand and complex compositional dynamics of the long-horizon tasks. We present Sequential Dexterity, a general system based on reinforcement learning (RL) that chains multiple dexterous policies for achieving long-horizon task goals. The core of the system is a transition feasibility function that progressively finetunes the sub-policies for enhancing chaining success rate, while also enables autonomous policy-switching for recovery from failures and bypassing redundant stages. Despite being trained only in simulation with a few task objects, our system demonstrates generalization capability to novel object shapes and is able to zero-shot transfer to a real-world robot equipped with a dexterous hand. Code and videos are available at https://sequential-dexterity.github.io
翻译:许多真实世界的操作任务由一系列彼此差异显著的子任务组成。此类长时域复杂任务凸显了灵巧手的潜力——其具备适应性与多功能性,可在无需重新抓取或外部工具的情况下无缝切换不同功能模式。然而,灵巧手的高维动作空间与长时域任务的复杂组合动力学特性带来了挑战。我们提出序列化灵巧性(Sequential Dexterity),一种基于强化学习(RL)的通用系统,通过链接多个灵巧策略实现长时域任务目标。该系统的核心是过渡可行性函数,其通过逐步微调子策略提升链接成功率,同时支持自主策略切换以从失败中恢复并绕过冗余阶段。尽管仅在仿真环境中使用少量任务物体进行训练,我们的系统仍展现出对未见物体形状的泛化能力,并可零样本迁移至配备灵巧手的真实机器人。代码与视频详见 https://sequential-dexterity.github.io