During task-oriented dialogues (TODs), human users naturally introduce chitchat that is beyond the immediate scope of the task, interfering with the flow of the conversation. To address this issue without the need for expensive manual data creation, we use few-shot prompting with Llama-2-70B to enhance the MultiWOZ dataset with user backstories, a typical example of chitchat interference in TODs. We assess the impact of this addition by testing two models: one trained solely on TODs and another trained on TODs with a preliminary chitchat interaction. Our analysis demonstrates that our enhanced dataset poses a challenge for these systems. Moreover, we demonstrate that our dataset can be effectively used for training purposes, enabling a system to consistently acknowledge the user's backstory while also successfully moving the task forward in the same turn, as confirmed by human evaluation. These findings highlight the benefits of generating novel chitchat-TOD scenarios to test TOD systems more thoroughly and improve their resilience to natural user interferences
翻译:在任务型对话(TOD)过程中,人类用户自然引入超出任务即时范围的闲聊,从而干扰对话流程。为解决这一问题,避免昂贵的人工数据创建,我们利用Llama-2-70B模型的少样本提示方法,为MultiWOZ数据集增强用户背景故事——这是TOD中闲聊干扰的典型示例。通过测试两种模型(一种仅基于TOD训练,另一种在TOD基础上增加初步闲聊交互训练),我们评估了数据增强的影响。分析表明,我们增强后的数据集对这些系统构成了挑战。此外,我们证明该数据集可有效用于训练目的,使系统能够始终如一地回应用户背景故事,同时在同一轮对话中顺利推进任务——这一结果已通过人工评估确认。这些发现凸显了生成新型闲聊-TOD场景的优势,有助于更全面地测试TOD系统,并提升其对自然用户干扰的鲁棒性。