In real-life conversations, the content is diverse, and there exists the one-to-many problem that requires diverse generation. Previous studies attempted to introduce discrete or Gaussian-based continuous latent variables to address the one-to-many problem, but the diversity is limited. Recently, diffusion models have made breakthroughs in computer vision, and some attempts have been made in natural language processing. In this paper, we propose DiffusionDialog, a novel approach to enhance the diversity of dialogue generation with the help of diffusion model. In our approach, we introduce continuous latent variables into the diffusion model. The problem of using latent variables in the dialog task is how to build both an effective prior of the latent space and an inferring process to obtain the proper latent given the context. By combining the encoder and latent-based diffusion model, we encode the response's latent representation in a continuous space as the prior, instead of fixed Gaussian distribution or simply discrete ones. We then infer the latent by denoising step by step with the diffusion model. The experimental results show that our model greatly enhances the diversity of dialog responses while maintaining coherence. Furthermore, in further analysis, we find that our diffusion model achieves high inference efficiency, which is the main challenge of applying diffusion models in natural language processing.
翻译:在真实对话中,内容具有多样性,存在需要多样化生成的“一对多”问题。以往研究尝试引入离散或基于高斯函数的连续潜变量来解决该问题,但多样性有限。近年来,扩散模型在计算机视觉领域取得突破,并已在自然语言处理中有所尝试。本文提出DiffusionDialog——一种借助扩散模型提升对话生成多样性的新方法。在该方法中,我们将连续潜变量引入扩散模型。对话任务中使用潜变量的核心问题在于:如何构建有效的潜空间先验,以及如何通过推理过程根据上下文获取合适的潜变量。通过结合编码器与基于潜变量的扩散模型,我们将响应的潜表示编码为连续空间中的先验,而非固定高斯分布或简单离散变量。随后,我们利用扩散模型逐步去噪以推断潜变量。实验结果表明,本模型在保持对话连贯性的同时显著增强了响应多样性。进一步分析发现,我们的扩散模型具有高推理效率,而这正是将扩散模型应用于自然语言处理的主要挑战。