Recent work has considered trust-aware decision making for human-robot collaboration (HRC) with a focus on model learning. In this paper, we are interested in enabling the HRC system to complete complex tasks specified using temporal logic that involve human trust. Since human trust in robots is not observable, we adopt the widely used partially observable Markov decision process (POMDP) framework for modelling the interactions between humans and robots. To specify the desired behaviour, we propose to use syntactically co-safe linear distribution temporal logic (scLDTL), a logic that is defined over predicates of states as well as belief states of partially observable systems. The incorporation of belief predicates in scLDTL enhances its expressiveness while simultaneously introducing added complexity. This also presents a new challenge as the belief predicates must be evaluated over the continuous (infinite) belief space. To address this challenge, we present an algorithm for solving the optimal policy synthesis problem. First, we enhance the belief MDP (derived by reformulating the POMDP) with a probabilistic labelling function. Then a product belief MDP is constructed between the probabilistically labelled belief MDP and the automaton translation of the scLDTL formula. Finally, we show that the optimal policy can be obtained by leveraging existing point-based value iteration algorithms with essential modifications. Human subject experiments with 21 participants on a driving simulator demonstrate the effectiveness of the proposed approach.
翻译:近期研究关注了人机协同中的信任感知决策问题,重点聚焦于模型学习。本文旨在使协作系统能够完成涉及人类信任的、用时序逻辑描述的复杂任务。由于人对机器人的信任不可直接观测,我们采用广泛使用的部分可观测马尔可夫决策过程(POMDP)框架来建模人机交互。为规范期望行为,我们提出使用语法协安全线性分布时序逻辑(scLDTL),该逻辑可定义于部分可观测系统的状态谓词及信念状态谓词之上。scLDTL中信念谓词的引入增强了其表达能力,但同时也增加了复杂度——这带来了新挑战:必须连续(无限)信念空间上评估信念谓词。为此,我们提出了求解最优策略综合问题的算法。首先,通过概率标记函数增强信念MDP(由POMDP重构得到);其次,在概率标记的信念MDP与scLDTL公式的自动机翻译之间构建乘积信念MDP;最后,证明可通过改进现有基于点的值迭代算法获得最优策略。基于驾驶模拟器的21名受试者实验验证了所提方法的有效性。