Human impressions of robot performance are often measured through surveys. As a more scalable and cost-effective alternative, we study the possibility of predicting people's impressions of robot behavior using non-verbal behavioral cues and machine learning techniques. To this end, we first contribute the SEAN TOGETHER Dataset consisting of observations of an interaction between a person and a mobile robot in a Virtual Reality simulation, together with impressions of robot performance provided by users on a 5-point scale. Second, we contribute analyses of how well humans and supervised learning techniques can predict perceived robot performance based on different combinations of observation types (e.g., facial, spatial, and map features). Our results show that facial expressions alone provide useful information about human impressions of robot performance; but in the navigation scenarios we tested, spatial features are the most critical piece of information for this inference task. Also, when evaluating results as binary classification (rather than multiclass classification), the F1-Score of human predictions and machine learning models more than doubles, showing that both are better at telling the directionality of robot performance than predicting exact performance ratings. Based on our findings, we provide guidelines for implementing these predictions models in real-world navigation scenarios.
翻译:人们通常通过问卷调查来衡量对机器人表现的主观印象。作为一种更具可扩展性和成本效益的替代方案,我们研究了利用非语言行为线索和机器学习技术预测人们对机器人行为印象的可能性。为此,我们首先贡献了SEAN TOGETHER数据集,该数据集包含虚拟现实模拟中人机交互的观测记录,以及用户对机器人表现按5分制提供的印象评分。其次,我们分析了人类与监督学习技术在不同观测类型组合(如面部特征、空间特征和地图特征)下预测感知机器人表现的效果。结果表明:面部表情本身即可提供关于用户对机器人表现印象的有效信息;但在我们测试的导航场景中,空间特征是该推理任务最关键的判断依据。此外,当以二分类(而非多分类)评估结果时,人类预测与机器学习模型的F1分数提升超过两倍,表明两者在判断机器人表现方向性上的能力显著优于精确预测评分。基于研究结论,我们为在实际导航场景中部署此类预测模型提供了指导原则。