Longitudinal Dialogues (LD) are the most challenging type of conversation for human-machine dialogue systems. LDs include the recollections of events, personal thoughts, and emotions specific to each individual in a sparse sequence of dialogue sessions. Dialogue systems designed for LDs should uniquely interact with the users over multiple sessions and long periods of time (e.g. weeks), and engage them in personal dialogues to elaborate on their feelings, thoughts, and real-life events. In this paper, we study the task of response generation in LDs. We evaluate whether general-purpose Pre-trained Language Models (PLM) are appropriate for this purpose. We fine-tune two PLMs, GePpeTto (GPT-2) and iT5, using a dataset of LDs. We experiment with different representations of the personal knowledge extracted from LDs for grounded response generation, including the graph representation of the mentioned events and participants. We evaluate the performance of the models via automatic metrics and the contribution of the knowledge via the Integrated Gradients technique. We categorize the natural language generation errors via human evaluations of contextualization, appropriateness and engagement of the user.
翻译:纵向对话(LD)是人机对话系统最具挑战性的对话类型。LD包含稀疏对话会话序列中每个个体特有的回忆、个人想法和情感。为LD设计的对话系统需在多会话及长时间跨度(如数周)内与用户进行独特交互,并通过个性化对话引导用户深入表达其感受、想法及现实生活事件。本文研究LD场景下的回复生成任务,评估通用预训练语言模型(PLM)的适用性。我们采用LD数据集对GePpeTto(GPT-2)和iT5两种PLM进行微调,实验性地探索从LD中提取个人知识的不同表示方式(包括提及事件及参与者的图表示)用于基于事实的回复生成。通过自动评估指标衡量模型性能,并基于积分梯度技术分析知识贡献度。结合上下文关联性、适切性和用户参与度的人工评估对自然语言生成错误进行分类。