Recent advancements in conversational systems have significantly enhanced human-machine interactions across various domains. However, training these systems is challenging due to the scarcity of specialized dialogue data. Traditionally, conversational datasets were created through crowdsourcing, but this method has proven costly, limited in scale, and labor-intensive. As a solution, the development of synthetic dialogue data has emerged, utilizing techniques to augment existing datasets or convert textual resources into conversational formats, providing a more efficient and scalable approach to dataset creation. In this survey, we offer a systematic and comprehensive review of multi-turn conversational data generation, focusing on three types of dialogue systems: open domain, task-oriented, and information-seeking. We categorize the existing research based on key components like seed data creation, utterance generation, and quality filtering methods, and introduce a general framework that outlines the main principles of conversation data generation systems. Additionally, we examine the evaluation metrics and methods for assessing synthetic conversational data, address current challenges in the field, and explore potential directions for future research. Our goal is to accelerate progress for researchers and practitioners by presenting an overview of state-of-the-art methods and highlighting opportunities to further research in this area.
翻译:近年来,对话系统的进步显著增强了各领域的人机交互能力。然而,由于专门对话数据的稀缺性,训练这些系统仍面临挑战。传统上,对话数据集通过众包方式创建,但该方法已被证明成本高昂、规模有限且劳动密集。作为解决方案,合成对话数据的开发应运而生,通过利用现有数据集增强技术或将文本资源转化为对话格式,为数据集创建提供了更高效且可扩展的方法。本综述对多轮对话数据生成进行了系统而全面的梳理,重点关注三类对话系统:开放域对话系统、任务导向型对话系统及信息寻求型对话系统。我们基于种子数据创建、话语生成与质量过滤方法等关键组件对现有研究进行分类,并提出了一个概括对话数据生成系统核心原理的通用框架。此外,我们探讨了评估合成对话数据的指标与方法,分析了该领域当前面临的挑战,并展望了未来潜在的研究方向。通过概述最新方法并强调可进一步开展研究的机遇,我们旨在加速研究人员与从业者的进展。