In this paper, we focus on inferring whether the given user command is clear, ambiguous, or infeasible in the context of interactive robotic agents utilizing large language models (LLMs). To tackle this problem, we first present an uncertainty estimation method for LLMs to classify whether the command is certain (i.e., clear) or not (i.e., ambiguous or infeasible). Once the command is classified as uncertain, we further distinguish it between ambiguous or infeasible commands leveraging LLMs with situational aware context in a zero-shot manner. For ambiguous commands, we disambiguate the command by interacting with users via question generation with LLMs. We believe that proper recognition of the given commands could lead to a decrease in malfunction and undesired actions of the robot, enhancing the reliability of interactive robot agents. We present a dataset for robotic situational awareness, consisting pair of high-level commands, scene descriptions, and labels of command type (i.e., clear, ambiguous, or infeasible). We validate the proposed method on the collected dataset, pick-and-place tabletop simulation. Finally, we demonstrate the proposed approach in real-world human-robot interaction experiments, i.e., handover scenarios.
翻译:本文聚焦于在利用大语言模型(LLMs)的交互式机器人场景中,推断给定用户指令是明确的、存在歧义的还是不可行的。为解决此问题,我们首先提出一种针对LLMs的不确定性估计方法,用于分类指令是否确定(即明确)或不确定(即存在歧义或不可行)。当指令被归类为不确定时,我们进一步利用具备情境感知上下文的LLMs,以零样本方式区分其为歧义指令还是不可行指令。对于歧义指令,我们通过LLMs生成问题与用户交互来消除歧义。我们相信,恰当识别给定指令有助于减少机器人的故障和不当行为,从而提升交互式机器人代理的可靠性。我们构建了一个面向机器人情境感知的数据集,包含高层指令、场景描述及指令类型标签(即明确、歧义或不可行)的配对数据。我们在收集的数据集及桌面拾放仿真场景上验证了所提方法。最后,我们在真实世界的人机交互实验(即交接场景)中演示了所提方法。