Recent advances in reinforcement learning (RL) have increased the promise of introducing cognitive assistance and automation to robot-assisted laparoscopic surgery (RALS). However, progress in algorithms and methods depends on the availability of standardized learning environments that represent skills relevant to RALS. We present LapGym, a framework for building RL environments for RALS that models the challenges posed by surgical tasks, and sofa_env, a diverse suite of 12 environments. Motivated by surgical training, these environments are organized into 4 tracks: Spatial Reasoning, Deformable Object Manipulation & Grasping, Dissection, and Thread Manipulation. Each environment is highly parametrizable for increasing difficulty, resulting in a high performance ceiling for new algorithms. We use Proximal Policy Optimization (PPO) to establish a baseline for model-free RL algorithms, investigating the effect of several environment parameters on task difficulty. Finally, we show that many environments and parameter configurations reflect well-known, open problems in RL research, allowing researchers to continue exploring these fundamental problems in a surgical context. We aim to provide a challenging, standard environment suite for further development of RL for RALS, ultimately helping to realize the full potential of cognitive surgical robotics. LapGym is publicly accessible through GitHub (https://github.com/ScheiklP/lap_gym).
翻译:近年来强化学习(RL)的进展增强了将认知辅助与自动化引入机器人辅助腹腔镜手术(RALS)的可行性。然而,算法与方法的进步取决于能否获得代表RALS相关技能的标准化学习环境。我们提出了LapGym,一个用于构建RALS的RL环境的框架,它模拟了手术任务带来的挑战;同时提出sofa_env,一套包含12个多样化环境的综合环境库。受外科训练启发,这些环境被组织为4个轨道:空间推理、可变形物体操作与抓取、解剖切割及线材操控。每个环境均具有高度参数化特性以增加难度,从而为新算法提供了高绩效上限。我们使用近端策略优化(PPO)建立了无模型RL算法的基线,探究了多种环境参数对任务难度的影响。最后,我们证明许多环境与参数配置反映了RL研究中广为人知的开放性问题,使研究者能在手术情境下继续探索这些基础课题。我们旨在为RALS中RL的进一步发展提供一套具有挑战性的标准化环境库,最终助力实现认知手术机器人的全部潜力。LapGym通过GitHub(https://github.com/ScheiklP/lap_gym)公开获取。