Large Language Models (LLMs) have demonstrated their ability to replicate human behaviors across a wide range of scenarios. However, their capability in handling complex, multi-character social interactions has yet to be fully explored, primarily due to the absence of robust, quantitative evaluation methods. This gap has slowed the development of agents proficient in more nuanced interactions beyond simple exchanges, for example, small talk. To address this challenge, we introduce the Multi-Agent Interaction Evaluation Framework (AntEval), encompassing a novel interaction framework and evaluation methods. The interaction framework aims to foster an complex interaction environment that bolsters information exchange and intention expression within social interactions. Furthermore, we introduce evaluation methods, including two metrics: Information Exchanging Precision (IEP) and Interaction Expressiveness Gap (IEG), designed for the quantitative and objective assessment of agents' interaction competencies. Our findings highlight the utility of these evaluative methods and show significant potential for improving LLMs' ability to construct agents that interact in a more natural manner with human-like intricacy.
翻译:大语言模型(LLMs)已展现出在多种场景中复现人类行为的能力。然而,它们在处理复杂、多角色社交互动方面的能力仍待充分探究,主要原因在于缺乏稳健的量化评估方法。这一空白阻碍了能够进行超越简单对话(例如闲聊)的微妙交互的智能体开发进程。为应对这一挑战,我们提出多智能体互动评估框架(AntEval),该框架包含新型互动架构与评估方法。互动架构旨在构建促进信息交换与意图表达的复杂交互环境,从而增强社交互动效果。在此基础上,我们引入评估方法,包含两个量化指标:信息交换精度(IEP)与互动表现力差距(IEG),用于客观量化评估智能体的交互能力。研究结果验证了这些评估方法的实用性,并揭示了提升LLMs构建更具自然交互能力与人类复杂性的智能体的显著潜力。