Many real-world applications of language models (LMs), such as writing assistance and code autocomplete, involve human-LM interaction. However, most benchmarks are non-interactive in that a model produces output without human involvement. To evaluate human-LM interaction, we develop a new framework, Human-AI Language-based Interaction Evaluation (HALIE), that defines the components of interactive systems and dimensions to consider when designing evaluation metrics. Compared to standard, non-interactive evaluation, HALIE captures (i) the interactive process, not only the final output; (ii) the first-person subjective experience, not just a third-party assessment; and (iii) notions of preference beyond quality (e.g., enjoyment and ownership). We then design five tasks to cover different forms of interaction: social dialogue, question answering, crossword puzzles, summarization, and metaphor generation. With four state-of-the-art LMs (three variants of OpenAI's GPT-3 and AI21 Labs' Jurassic-1), we find that better non-interactive performance does not always translate to better human-LM interaction. In particular, we highlight three cases where the results from non-interactive and interactive metrics diverge and underscore the importance of human-LM interaction for LM evaluation.
翻译:语言模型(LM)在写作辅助和代码自动补全等实际应用场景中常常涉及人机交互。然而,现有大多数基准测试采用非交互式模式,即模型在无人类参与下生成输出。为评估人类-LM交互,我们开发了新框架——人类-人工智能语言交互评估(HALIE),该框架定义了交互系统的组成要素及设计评估指标时需考量的维度。相较于标准非交互式评估,HALIE能捕捉:(i) 交互过程(不仅限于最终输出);(ii) 第一人称主观体验(而非第三方评估);(iii) 超越质量的偏好概念(如愉悦感与所有权)。我们进一步设计了覆盖不同交互形式的五项任务:社交对话、问答、填字游戏、文本摘要及隐喻生成。基于四个前沿LM(OpenAI GPT-3的三个变体及AI21 Labs的Jurassic-1)的实验表明,非交互式表现更优并不总能转化为更佳的人机交互效果。我们重点指出三组非交互式与交互式指标结果相背离的案例,强调人机交互评估对LM评测的重要性。