Robot learning methods have recently made great strides, but generalization and robustness challenges still hinder their widespread deployment. Failing to detect and address potential failures renders state-of-the-art learning systems not combat-ready for high-stakes tasks. Recent advances in interactive imitation learning have presented a promising framework for human-robot teaming, enabling the robots to operate safely and continually improve their performances over long-term deployments. Nonetheless, existing methods typically require constant human supervision and preemptive feedback, limiting their practicality in realistic domains. This work aims to endow a robot with the ability to monitor and detect errors during task execution. We introduce a model-based runtime monitoring algorithm that learns from deployment data to detect system anomalies and anticipate failures. Unlike prior work that cannot foresee future failures or requires failure experiences for training, our method learns a latent-space dynamics model and a failure classifier, enabling our method to simulate future action outcomes and detect out-of-distribution and high-risk states preemptively. We train our method within an interactive imitation learning framework, where it continually updates the model from the experiences of the human-robot team collected using trustworthy deployments. Consequently, our method reduces the human workload needed over time while ensuring reliable task execution. Our method outperforms the baselines across system-level and unit-test metrics, with 23% and 40% higher success rates in simulation and on physical hardware, respectively. More information at https://ut-austin-rpl.github.io/sirius-runtime-monitor/
翻译:近年来,机器人学习方法取得了显著进展,但泛化能力和鲁棒性挑战仍阻碍其广泛应用。未能检测和应对潜在故障导致最先进的学习系统无法胜任高风险任务。交互式模仿学习的最新进展为人机协作提供了有前景的框架,使机器人能够在长期部署中安全运行并持续改进性能。然而,现有方法通常需要持续的人类监督和先发制人的反馈,限制了其在现实场景中的实用性。本研究旨在赋予机器人在任务执行过程中监测和检测错误的能力。我们提出一种基于模型的运行时监控算法,该算法从部署数据中学习,以检测系统异常并预测故障。与无法预判未来故障或需要故障经验进行训练的现有方法不同,我们的方法学习潜在空间动力学模型和故障分类器,从而能够模拟未来动作结果并主动检测分布外状态和高风险状态。我们将其嵌入交互式模仿学习框架进行训练,通过可信部署收集人机协作经验,持续更新模型。因此,我们的方法在确保可靠任务执行的同时,逐步降低所需的人类工作量。在系统级和单元测试指标上,我们的方法均优于基线模型,在仿真环境和真实硬件上的成功率分别提高了23%和40%。更多信息请访问 https://ut-austin-rpl.github.io/sirius-runtime-monitor/