Robots assist humans in various activities, from daily living public service (e.g., airports and restaurants), and to collaborative manufacturing. However, it is risky to assume that the knowledge and strategies robots learned from one group of people can apply to other groups. The discriminatory performance of robots will undermine their service quality for some people, ignore their service requests, and even offend them. Therefore, it is critically important to mitigate bias in robot decision-making for more fair services. In this paper, we designed a self-reflective mechanism -- Fairness-Sensitive Policy Gradient Reinforcement Learning (FSPGRL), to help robots to self-identify biased behaviors during interactions with humans. FSPGRL identifies bias by examining the abnormal update along particular gradients and updates the policy network to support fair decision-making for robots. To validate FSPGRL's effectiveness, a human-centered service scenario, "A robot is serving people in a restaurant," was designed. A user study was conducted; 24 human subjects participated in generating 1,000 service demonstrations. Four commonly-seen issues "Willingness Issue," "Priority Issue," "Quality Issue," "Risk Issue" were observed from robot behaviors. By using FSPGRL to improve robot decisions, robots were proven to have a self-bias detection capability for a more fair service. We have achieved the suppression of bias and improved the quality during the process of robot learning to realize a relatively fair model.
翻译:机器人在各种活动中协助人类,从日常生活公共服务(例如机场和餐厅)到协作制造。然而,假设机器人从一组人群中学到的知识和策略可以应用于其他人群是有风险的。机器人的歧视性表现会削弱其对某些人的服务质量,忽视其服务请求,甚至冒犯他们。因此,减轻机器人决策中的偏见以提供更公平的服务至关重要。本文设计了一种自反机制——面向公平性的策略梯度强化学习(FSPGRL),以帮助机器人在与人类互动过程中自我识别偏见行为。FSPGRL通过检查特定梯度上的异常更新来识别偏见,并更新策略网络以支持机器人做出公平决策。为验证FSPGRL的有效性,我们设计了一个以人为中心的服务场景“餐厅中为顾客服务的机器人”,并开展了一项用户研究,24名受试者参与生成了1000次服务演示。从机器人行为中观察到了四种常见问题:“意愿问题”、“优先级问题”、“质量问题”和“风险问题”。通过使用FSPGRL改进机器人决策,机器人被证明具有自我偏见检测能力,从而提供更公平的服务。我们实现了在机器人学习过程中抑制偏见并提升质量,最终获得了一个相对公平的模型。