Norms are an important component of the social fabric of society by prescribing expected behaviour. In Multi-Agent Systems (MAS), agents interacting within a society are equipped to possess social capabilities such as reasoning about norms and trust. Norms have long been of interest within the Normative Multi-Agent Systems community with researchers studying topics such as norm emergence, norm violation detection and sanctioning. However, these studies have some limitations: they are often limited to simple domains, norms have been represented using a variety of representations with no standard approach emerging, and the symbolic reasoning mechanisms generally used may suffer from a lack of extensibility and robustness. In contrast, Large Language Models (LLMs) offer opportunities to discover and reason about norms across a large range of social situations. This paper evaluates the capability of LLMs to detecting norm violations. Based on simulated data from 80 stories in a household context, with varying complexities, we investigated whether 10 norms are violated. For our evaluations we first obtained the ground truth from three human evaluators for each story. Then, the majority result was compared against the results from three well-known LLM models (Llama 2 7B, Mixtral 7B and ChatGPT-4). Our results show the promise of ChatGPT-4 for detecting norm violations, with Mixtral some distance behind. Also, we identify areas where these models perform poorly and discuss implications for future work.
翻译:规范通过规定期望行为,成为构成社会结构的重要元素。在多智能体系统中,交互的智能体具备诸如规范推理和信任等社会能力。规范性多智能体系统社区长期以来对规范相关研究抱有浓厚兴趣,学者们探讨了规范涌现、规范违反检测与制裁等主题。然而,现有研究存在若干局限:通常局限于简单领域,规范表征方式多样且缺乏统一标准,所采用的符号推理机制在可扩展性与鲁棒性方面存在不足。相比之下,大语言模型为跨广泛社会情境发现并推理规范提供了新的可能。本文评估了大语言模型检测规范违反的能力。我们基于80个家庭场景故事生成的仿真数据(包含不同复杂度),研究其中10项规范是否被违反。评估过程中,首先获取三位人类评估者对每个故事的真值标注,随后将多数投票结果与三个知名大语言模型(Llama 2 7B、Mixtral 7B与ChatGPT-4)的输出进行比较。结果表明,ChatGPT-4在规范违反检测方面表现优异,Mixtral则明显落后。同时,本文识别了这些模型的薄弱环节,并讨论了未来工作的启示。