Automating robotic surgery via learning from demonstration (LfD) techniques is extremely challenging. This is because surgical tasks often involve sequential decision-making processes with complex interactions of physical objects and have low tolerance for mistakes. Prior works assume that all demonstrations are fully observable and optimal, which might not be practical in the real world. This paper introduces a sample-efficient method that learns a robust reward function from a limited amount of ranked suboptimal demonstrations consisting of partial-view point cloud observations. The method then learns a policy by optimizing the learned reward function using reinforcement learning (RL). We show that using a learned reward function to obtain a policy is more robust than pure imitation learning. We apply our approach on a physical surgical electrocautery task and demonstrate that our method can perform well even when the provided demonstrations are suboptimal and the observations are high-dimensional point clouds. Code and videos available here: https://sites.google.com/view/lfdinelectrocautery
翻译:通过从演示中学习(LfD)技术实现机器人手术自动化极具挑战性。这是因为手术任务通常涉及具有复杂物理对象交互的序贯决策过程,且对错误的容忍度极低。先前研究假设所有演示均可完全观测且最优,这在现实世界中可能并不实际。本文提出一种样本高效方法,能从少量包含部分视角点云观测的排序次优演示中学习鲁棒奖励函数。该方法随后通过强化学习(RL)优化所学习的奖励函数来推导策略。研究表明,利用学习到的奖励函数获取策略比纯模仿学习更具鲁棒性。我们将该方法应用于物理手术电外科任务,并证明即便提供的演示为次优且观测为高维点云时,该方法仍能表现良好。代码与视频详见:https://sites.google.com/view/lfdinelectrocautery