Open-domain dialogue system usually requires different sources of knowledge to generate more informative and evidential responses. However, existing knowledge-grounded dialogue systems either focus on a single knowledge source or overlook the dependency between multiple sources of knowledge, which may result in generating inconsistent or even paradoxical responses. To incorporate multiple knowledge sources and dependencies between them, we propose SAFARI, a novel framework that leverages the exceptional capabilities of large language models (LLMs) in planning, understanding, and incorporating under both supervised and unsupervised settings. Specifically, SAFARI decouples the knowledge grounding into multiple sources and response generation, which allows easy extension to various knowledge sources including the possibility of not using any sources. To study the problem, we construct a personalized knowledge-grounded dialogue dataset \textit{\textbf{K}nowledge \textbf{B}ehind \textbf{P}ersona}~(\textbf{KBP}), which is the first to consider the dependency between persona and implicit knowledge. Experimental results on the KBP dataset demonstrate that the SAFARI framework can effectively produce persona-consistent and knowledge-enhanced responses.
翻译:开放域对话系统通常需要多种知识源以生成更具信息性和证据性的回复。然而,现有基于知识的对话系统要么聚焦于单一知识源,要么忽略多源知识之间的依赖性,这可能导致生成不一致甚至矛盾的回复。为整合多源知识及其依赖关系,我们提出SAFARI框架——一种创新性方法,利用大语言模型在规划、理解及融合方面的卓越能力,同时支持监督与无监督设置。具体而言,SAFARI将知识融合过程解耦为多源知识规划与回复生成两个阶段,从而易于扩展至多种知识源(包括不使用任何知识源)的情况。针对该研究问题,我们构建了个性化知识对话数据集《人格中的隐式知识》(KBP),该数据集首次考虑人格特征与隐式知识之间的依赖关系。在KBP数据集上的实验结果表明,SAFARI框架能够有效生成保持人格一致性且富含知识增强的回复。