Generalizable manipulation skills, which can be composed to tackle long-horizon and complex daily chores, are one of the cornerstones of Embodied AI. However, existing benchmarks, mostly composed of a suite of simulatable environments, are insufficient to push cutting-edge research works because they lack object-level topological and geometric variations, are not based on fully dynamic simulation, or are short of native support for multiple types of manipulation tasks. To this end, we present ManiSkill2, the next generation of the SAPIEN ManiSkill benchmark, to address critical pain points often encountered by researchers when using benchmarks for generalizable manipulation skills. ManiSkill2 includes 20 manipulation task families with 2000+ object models and 4M+ demonstration frames, which cover stationary/mobile-base, single/dual-arm, and rigid/soft-body manipulation tasks with 2D/3D-input data simulated by fully dynamic engines. It defines a unified interface and evaluation protocol to support a wide range of algorithms (e.g., classic sense-plan-act, RL, IL), visual observations (point cloud, RGBD), and controllers (e.g., action type and parameterization). Moreover, it empowers fast visual input learning algorithms so that a CNN-based policy can collect samples at about 2000 FPS with 1 GPU and 16 processes on a regular workstation. It implements a render server infrastructure to allow sharing rendering resources across all environments, thereby significantly reducing memory usage. We open-source all codes of our benchmark (simulator, environments, and baselines) and host an online challenge open to interdisciplinary researchers.
翻译:通用化操作技能是具身智能的基石之一,它们能够被组合以解决长时域和复杂的日常家务。然而,现有基准大多由一组可模拟环境构成,由于缺乏物体层级的拓扑与几何变化、未基于全动态仿真,或缺乏对多种操作任务的天然支持,不足以推动前沿研究。为此,我们提出下一代SAPIEN ManiSkill基准——ManiSkill2,以解决研究人员在使用基准进行通用化操作技能研究时经常遇到的关键痛点。ManiSkill2包含20个操作任务族,拥有2000多个物体模型和400多万帧演示数据,覆盖固定/移动基座、单/双臂、刚体/软体操作任务,并由全动态引擎模拟提供2D/3D输入数据。它定义了统一的接口与评估协议,以支持多种算法(如经典的感知-规划-执行、强化学习、模仿学习)、视觉观测(点云、RGBD)和控制器(如动作类型与参数化)。此外,它加速了快速视觉输入学习算法,使得基于CNN的策略可在普通工作站上以约2000 FPS的速度收集样本(1 GPU、16进程)。为实现渲染资源的跨环境共享,我们构建了渲染服务器基础设施,从而显著降低内存使用。我们开源了基准的全部代码(仿真器、环境和基线),并面向跨学科研究人员举办在线挑战赛。