Protecting against adversarial attacks is a common multiagent problem. Attackers in the real world are predominantly human actors, and the protection methods often incorporate opponent models to improve the performance when facing humans. Previous results show that modeling human behavior can significantly improve the performance of the algorithms. However, modeling humans correctly is a complex problem, and the models are often simplified and assume humans make mistakes according to some distribution or train parameters for the whole population from which they sample. In this work, we use data gathered by psychologists who identified personality types that increase the likelihood of performing malicious acts. However, in the previous work, the tests on a handmade game could not show strategic differences between the models. We created a novel model that links its parameters to psychological traits. We optimized over parametrized games and created games in which the differences are profound. Our work can help with automatic game generation when we need a game in which some models will behave differently and to identify situations in which the models do not align.
翻译:针对对抗性攻击的防御是一个常见的多智能体问题。现实世界中的攻击者主要是人类行为主体,防御方法通常融入对手模型以提升面对人类时的性能。先前研究结果表明,对人类行为进行建模可以显著提升算法性能。然而,正确建模人类是一个复杂问题,现有模型往往经过简化处理,或假设人类依据某种分布产生错误,或针对整个人群采样训练参数。在本研究中,我们使用了心理学家收集的数据,这些数据识别出会增加恶意行为发生概率的人格类型。但在先前工作中,基于手工设计的博弈进行的测试无法展现模型间的策略差异。我们创建了一种新型模型,将其参数与心理特质相关联。通过对参数化博弈进行优化,我们构建了模型差异显著的博弈场景。本研究有助于在需要不同模型呈现差异化行为的情景中实现自动博弈生成,并识别模型间不一致的情况。