Applications that could benefit from automatic understanding of human-human conversations often come with challenges associated with private information in real-world data such as call center or clinical conversations. Working with protected data also increases costs of annotation, which limits technology development. To address these challenges, we propose DIALGEN, a human-in-the-loop semi-automated dialogue generation framework. DIALGEN uses a language model (ChatGPT) that can follow schema and style specifications to produce fluent conversational text, generating a complex conversation through iteratively generating subdialogues and using human feedback to correct inconsistencies or redirect the flow. In experiments on structured summarization of agent-client information gathering calls, framed as dialogue state tracking, we show that DIALGEN data enables significant improvement in model performance.
翻译:能够从自动理解人人对话中受益的应用,往往面临真实世界数据(如呼叫中心或临床对话)中隐私信息相关的挑战。处理受保护数据还会增加标注成本,从而限制了技术发展。为应对这些挑战,我们提出DIALGEN——一种人机协作的半自动化对话生成框架。DIALGEN利用可遵循架构和风格规范的语言模型(ChatGPT)生成流畅的对话文本,通过迭代生成子对话并利用人类反馈纠正不一致或引导对话流向,从而构建复杂对话。在面向对话状态追踪框架下的代理人-客户信息收集通话结构化摘要实验中,我们证明DIALGEN数据能够显著提升模型性能。