One of the major impediments to the development of new task-oriented dialogue (TOD) systems is the need for human evaluation at multiple stages and iterations of the development process. In an effort to move toward automated evaluation of TOD, we propose a novel user simulator built using recently developed large pretrained language models (LLMs). In order to increase the linguistic diversity of our system relative to the related previous work, we do not fine-tune the LLMs used by our system on existing TOD datasets; rather we use in-context learning to prompt the LLMs to generate robust and linguistically diverse output with the goal of simulating the behavior of human interlocutors. Unlike previous work, which sought to maximize goal success rate (GSR) as the primary metric of simulator performance, our goal is a system which achieves a GSR similar to that observed in human interactions with TOD systems. Using this approach, our current simulator is effectively able to interact with several TOD systems, especially on single-intent conversational goals, while generating lexically and syntactically diverse output relative to previous simulators that rely upon fine-tuned models. Finally, we collect a Human2Bot dataset of humans interacting with the same TOD systems with which we experimented in order to better quantify these achievements.
翻译:任务型对话系统开发的主要障碍之一是在开发过程的多个阶段和迭代中需要进行人工评估。为了实现任务型对话的自动化评估,我们提出了一种新型用户模拟器,该模拟器基于近期开发的大规模预训练语言模型构建。为了增加系统相对于相关先前工作的语言多样性,我们不对系统所使用的语言模型进行微调;相反,我们利用上下文学习来引导语言模型生成稳健且语言多样化的输出,以模拟人类对话者的行为。与先前工作将最大目标成功率作为模拟器性能主要指标不同,我们的目标是实现一个在与任务型对话系统交互时达到与人类交互相似目标的系统。采用这种方法,我们当前的模拟器能够有效地与多个任务型对话系统进行交互,尤其是在处理单意图对话目标时,同时生成相比依赖微调模型的先前模拟器具有更大词汇和句法多样性的输出。最后,我们收集了人类与相同任务型对话系统交互的Human2Bot数据集,以便更好地量化这些成果。