Over the past two decades, dialogue modeling has made significant strides, moving from simple rule-based responses to personalized and persuasive response generation. However, despite these advancements, the objective functions and evaluation metrics for dialogue generation have remained stagnant, i.e., cross-entropy and BLEU, respectively. These lexical-based metrics have the following key limitations: (a) word-to-word matching without semantic consideration: It assigns the same credit for failure to generate 'nice' and 'rice' for 'good'. (b) missing context attribute for evaluating the generated response: Even if a generated response is relevant to the ongoing dialogue context, it may still be penalized for not matching the gold utterance provided in the corpus. In this paper, we first investigate these limitations comprehensively and propose a new loss function called Semantic Infused Contextualized diaLogue (SemTextualLogue) loss function. Furthermore, we formulate a new evaluation metric called Dialuation, which incorporates both context relevance and semantic appropriateness while evaluating a generated response. We conducted experiments on two benchmark dialogue corpora, encompassing both task-oriented and open-domain scenarios. We found that the dialogue generation model trained with SemTextualLogue loss attained superior performance (in both quantitative and qualitative evaluation) compared to the traditional cross-entropy loss function across the datasets and evaluation metrics.
翻译:过去二十年间,对话建模取得了显著进展,从简单的基于规则的响应发展为个性化且具有说服力的响应生成。然而,尽管取得了这些进步,对话生成的目标函数和评估指标却一直停滞不前,即分别采用交叉熵和BLEU。这些基于词汇的指标存在以下关键局限性:(a)缺乏语义考量的词对词匹配:它将生成“nice”和“rice”以替代“good”的失败视为同等错误。(b)缺失评估生成响应的上下文属性:即使生成的响应与当前对话上下文相关,它仍可能因未匹配语料库中提供的标准话语而受到惩罚。在本文中,我们首先全面探讨了这些局限性,并提出了一种新的损失函数,称为语义融合上下文化对话损失函数(SemTextualLogue)。此外,我们制定了一种新的评估指标Dialuation,其在评估生成响应时同时纳入了上下文相关性和语义恰当性。我们在两个涵盖任务导向型和开放域场景的基准对话语料库上进行了实验。我们发现,使用SemTextualLogue损失函数训练的对话生成模型在数据集和评估指标上(无论在定量还是定性评估中)均优于传统交叉熵损失函数。