Text response generation for multimodal task-oriented dialog systems, which aims to generate the proper text response given the multimodal context, is an essential yet challenging task. Although existing efforts have achieved compelling success, they still suffer from two pivotal limitations: 1) overlook the benefit of generative pre-training, and 2) ignore the textual context related knowledge. To address these limitations, we propose a novel dual knowledge-enhanced generative pretrained language model for multimodal task-oriented dialog systems (DKMD), consisting of three key components: dual knowledge selection, dual knowledge-enhanced context learning, and knowledge-enhanced response generation. To be specific, the dual knowledge selection component aims to select the related knowledge according to both textual and visual modalities of the given context. Thereafter, the dual knowledge-enhanced context learning component targets seamlessly integrating the selected knowledge into the multimodal context learning from both global and local perspectives, where the cross-modal semantic relation is also explored. Moreover, the knowledge-enhanced response generation component comprises a revised BART decoder, where an additional dot-product knowledge-decoder attention sub-layer is introduced for explicitly utilizing the knowledge to advance the text response generation. Extensive experiments on a public dataset verify the superiority of the proposed DKMD over state-of-the-art competitors.
翻译:针对多模态任务型对话系统中的文本响应生成任务——即根据多模态上下文生成恰当的文本响应——是一项重要且具有挑战性的工作。尽管现有研究已取得显著进展,但仍存在两个关键局限:1)忽略了生成式预训练的优势,2)未能利用与文本上下文相关的知识。为解决上述问题,我们提出了一种新颖的基于双重知识增强生成式预训练语言模型的多模态任务型对话系统(DKMD),包含三个核心组件:双重知识选择、双重知识增强的上下文学习以及知识增强的响应生成。具体而言,双重知识选择组件旨在根据给定上下文的文本与视觉模态选择相关知识;随后,双重知识增强的上下文学习组件从全局与局部视角将所选知识无缝集成到多模态上下文学习中,并探索跨模态语义关联;此外,知识增强的响应生成组件包含改进的BART解码器,其中引入额外的点积知识-解码器注意力子层,通过显式利用知识提升文本响应生成。在公开数据集上的大量实验验证了所提出的DKMD方法相较于当前最优方法的优越性。