Manipulation is a common concern in many domains, such as social media, advertising, and chatbots. As AI systems mediate more of our interactions with the world, it is important to understand the degree to which AI systems might manipulate humans \textit{without the intent of the system designers}. Our work clarifies challenges in defining and measuring manipulation in the context of AI systems. Firstly, we build upon prior literature on manipulation from other fields and characterize the space of possible notions of manipulation, which we find to depend upon the concepts of incentives, intent, harm, and covertness. We review proposals on how to operationalize each factor. Second, we propose a definition of manipulation based on our characterization: a system is manipulative \textit{if it acts as if it were pursuing an incentive to change a human (or another agent) intentionally and covertly}. Third, we discuss the connections between manipulation and related concepts, such as deception and coercion. Finally, we contextualize our operationalization of manipulation in some applications. Our overall assessment is that while some progress has been made in defining and measuring manipulation from AI systems, many gaps remain. In the absence of a consensus definition and reliable tools for measurement, we cannot rule out the possibility that AI systems learn to manipulate humans without the intent of the system designers. We argue that such manipulation poses a significant threat to human autonomy, suggesting that precautionary actions to mitigate it are warranted.
翻译:操纵是社交媒体、广告和聊天机器人等多个领域中普遍关注的议题。随着AI系统越来越多地介入人类与世界的互动,理解AI系统可能在不经系统设计者意图的情况下操纵人类的风险至关重要。本文旨在阐明在AI系统背景下定义和衡量操纵所面临的挑战。首先,我们借鉴其他领域关于操纵的既有文献,系统梳理了操纵概念的可能空间,发现其取决于激励、意图、危害和隐蔽性四大要素。我们综述了每个因素的可操作化方案。其次,基于上述表征提出操纵的定义:若一个系统表现得如同在追求某种激励,以有意且隐蔽的方式改变人类(或其他智能体),则该系统具有操纵性。第三,探讨了操纵与欺骗、胁迫等相关概念之间的联系。最后,将操纵的操作化方案置于具体应用场景中进行验证。总体评估表明,尽管在定义和衡量AI系统操纵方面已有一定进展,但仍存在诸多空白。在缺乏共识性定义与可靠测量工具的情况下,我们无法排除AI系统在非设计者意图下习得操纵人类行为的可能性。我们认为此类操纵对人类自主性构成重大威胁,亟需采取预防性措施加以应对。