Knowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support knowledge. Even in the era of large language models, response generation grounded in knowledge retrieved from additional up-to-date sources remains a practically important approach. While prior work using single-source knowledge has shown a clear positive correlation between the performances of knowledge selection and response generation, there are no existing multi-source datasets for evaluating support knowledge retrieval. Further, prior work has assumed that the knowledge sources available at test time are the same as during training. This unrealistic assumption unnecessarily handicaps models, as new knowledge sources can become available after a model is trained. In this paper, we present a high-quality benchmark named multi-source Wizard of Wikipedia (Ms.WoW) for evaluating multi-source dialogue knowledge selection and response generation. Unlike existing datasets, it contains clean support knowledge, grounded at the utterance level and partitioned into multiple knowledge sources. We further propose a new challenge, dialogue knowledge plug-and-play, which aims to test an already trained dialogue model on using new support knowledge from previously unseen sources in a zero-shot fashion.
翻译:基于知识的开放域对话生成旨在构建能够利用挖掘出的支撑知识与人类进行聊天的闲聊系统。已有研究表明,多种类型和来源的知识均可作为有效的支撑知识。即便在大语言模型时代,基于从额外最新来源检索到的知识进行回复生成仍然是一种具有实际重要性的方法。虽然先前基于单一来源知识的研究已表明知识选择与回复生成性能之间存在明确的正相关关系,但目前尚不存在用于评估支撑知识检索的多来源数据集。此外,先前研究假设测试时可用的知识来源与训练时相同。这一不切实际的假设不必要地限制了模型的能力,因为训练完成后可能涌现新的知识来源。本文提出了名为多来源维基百科向导(Ms.WoW)的高质量基准,用于评估多来源对话知识选择与回复生成。与现有数据集不同,该基准包含经语轮级标注且划分为多个知识来源的干净支撑知识。我们进一步提出了对话知识即插即用的新挑战,旨在测试已训练的对话模型在零样本场景下使用来自先前未见来源的新支撑知识的能力。