Companies, organizations, and governments increasingly exploit Language Models' (LM) remarkable capability to display agent-like behavior. As LMs are adopted to perform tasks with growing autonomy, there exists an urgent need for reliable and scalable evaluation benchmarks. Current, predominantly static LM benchmarks are ill-suited to evaluate such dynamic applications. Thus, we propose jointly evaluating LM performance and alignment through the lenses of negotiation games. We argue that this common task better reflects real-world deployment conditions while offering insights into LMs' decision-making processes. Crucially, negotiation games allow us to study multi-turn, and cross-model interactions, modulate complexity, and side-step accidental data leakage in evaluation. We report results for six publicly accessible LMs from several major providers on a variety of negotiation games, evaluating both self-play and cross-play performance. Noteworthy findings include: (i) open-source models are currently unable to complete these tasks; (ii) cooperative bargaining games prove challenging; and (iii) the most powerful models do not always "win".
翻译:企业、组织及政府正日益利用语言模型展现类代理行为的卓越能力。随着语言模型被用于执行自主性日益增长的任务,当前迫切需要可靠且可扩展的评估基准。现有以静态评估为主的基准方法并不适用于此类动态应用的评估。为此,我们提出通过谈判博弈的视角联合评估语言模型的性能与对齐性。我们认为这一常见任务既能更好地反映实际部署条件,又能揭示语言模型的决策过程。关键的是,谈判博弈使我们能够研究多轮交互与跨模型交互、调节任务复杂度,并规避评估中偶然发生的数据泄露问题。我们针对来自多家主要供应商的六种公开可访问语言模型,在多种谈判博弈场景下报告了自博弈与交叉博弈的性能评估结果。值得关注的发现包括:(i) 开源模型目前无法完成这些任务;(ii) 合作性讨价还价博弈具有挑战性;(iii) 最强大的模型并不总能"获胜"。