Tasks where robots must cooperate with humans, such as navigating around a cluttered home or sorting everyday items, are challenging because they exhibit a wide range of valid actions that lead to similar outcomes. Moreover, zero-shot cooperation between human-robot partners is an especially challenging problem because it requires the robot to infer and adapt on the fly to a latent human intent, which could vary significantly from human to human. Recently, deep learned motion prediction models have shown promising results in predicting human intent but are prone to being confidently incorrect. In this work, we present Risk-Calibrated Interactive Planning (RCIP), which is a framework for measuring and calibrating risk associated with uncertain action selection in human-robot cooperation, with the fundamental idea that the robot should ask for human clarification when the risk associated with the uncertainty in the human's intent cannot be controlled. RCIP builds on the theory of set-valued risk calibration to provide a finite-sample statistical guarantee on the cumulative loss incurred by the robot while minimizing the cost of human clarification in complex multi-step settings. Our main insight is to frame the risk control problem as a sequence-level multi-hypothesis testing problem, allowing efficient calibration using a low-dimensional parameter that controls a pre-trained risk-aware policy. Experiments across a variety of simulated and real-world environments demonstrate RCIP's ability to predict and adapt to a diverse set of dynamic human intents.
翻译:机器人需要与人类协作完成的任务,例如在杂乱的家庭环境中导航或分类日常物品,具有很大挑战性,因为这些任务中存在大量导致相似结果的合法动作。此外,人机伙伴之间的零样本协作尤其棘手,因为机器人需要即时推断并适应潜在的人类意图,而不同个体的意图可能差异显著。近年来,深度学习的运动预测模型在预测人类意图方面展现出良好前景,但容易产生过度自信的错误判断。本研究提出风险校准交互规划(RCIP)框架,用于衡量和校准人机协作中不确定动作选择的相关风险,其核心理念是:当人类意图不确定性带来的风险无法控制时,机器人应主动寻求人类澄清。RCIP基于集合风险校准理论,在复杂多步交互场景中,能在最小化人类澄清成本的同时,为机器人累积损失提供有限样本统计保证。我们的主要洞察是将风险控制问题转化为序列级多重假设检验问题,通过控制预训练风险感知策略的低维参数实现高效校准。在多种模拟和真实环境中的实验证明,RCIP能够预测并适应多样化的动态人类意图。