With the growing capabilities and pervasiveness of AI systems, societies must collectively choose between reduced human autonomy, endangered democracies and limited human rights, and AI that is aligned to human and social values, nurturing collaboration, resilience, knowledge and ethical behaviour. In this chapter, we introduce the notion of self-reflective AI systems for meaningful human control over AI systems. Focusing on decision support systems, we propose a framework that integrates knowledge from psychology and philosophy with formal reasoning methods and machine learning approaches to create AI systems responsive to human values and social norms. We also propose a possible research approach to design and develop self-reflective capability in AI systems. Finally, we argue that self-reflective AI systems can lead to self-reflective hybrid systems (human + AI), thus increasing meaningful human control and empowering human moral reasoning by providing comprehensible information and insights on possible human moral blind spots.
翻译:随着人工智能系统能力的增强与普及,社会必须在减少人类自主性、危及民主制度、限制人权,与追求契合人类及社会价值观、促进协作、韧性、知识与道德行为的人工智能之间做出集体选择。本章提出"自我反思式人工智能系统"概念,以实现对人类有意义的人工智能系统控制。聚焦决策支持系统,我们构建了一个整合心理学与哲学知识、形式化推理方法及机器学习技术的框架,使人工智能系统能够响应人类价值观与社会规范。同时提出一种可能的研究路径,用于设计与开发人工智能系统的自我反思能力。最终论证自我反思式人工智能系统能够催生自我反思式混合系统(人类+人工智能),通过提供可理解的信息及对人类潜在道德盲点的洞察,增强人类有意义控制并强化道德推理能力。