The reinforcement learning algorithms have often been applied to social robots. However, most reinforcement learning algorithms were not optimized for the use of social robots, and consequently they may bore users. We proposed a new reinforcement learning method specialized for the social robot, the FRAC-Q-learning, that can avoid user boredom. The proposed algorithm consists of a forgetting process in addition to randomizing and categorizing processes. This study evaluated interest and boredom hardness scores of the FRAC-Q-learning by a comparison with the traditional Q-learning. The FRAC-Q-learning showed significantly higher trend of interest score, and indicated significantly harder to bore users compared to the traditional Q-learning. Therefore, the FRAC-Q-learning can contribute to develop a social robot that will not bore users. The proposed algorithm can also find applications in Web-based communication and educational systems. This paper presents the entire process, detailed implementation and a detailed evaluation method of the of the FRAC-Q-learning for the first time.
翻译:强化学习算法常被应用于社交机器人。然而,大多数强化学习算法并非针对社交机器人的使用场景进行优化,因此可能导致用户产生厌倦感。我们提出了一种专门针对社交机器人的新型强化学习方法——FRAC-Q-Learning,该方法能够避免用户厌倦。该算法在随机化与分类过程的基础上增加了一个遗忘过程。本研究通过与传统的Q-Learning进行对比,评估了FRAC-Q-Learning的兴趣分数与厌倦难度分数。相较于传统Q-Learning,FRAC-Q-Learning显著提升了兴趣分数趋势,并显示出显著更强的用户厌倦抵抗能力。因此,FRAC-Q-Learning有助于开发不会使用户感到厌倦的社交机器人。该算法还可应用于基于网络的通信与教育系统。本文首次完整呈现了FRAC-Q-Learning的整体流程、详细实现方法及系统的评估方案。