Task-oriented dialogue systems often assist users with personal or confidential matters. For this reason, the developers of such a system are generally prohibited from observing actual usage. So how can they know where the system is failing and needs more training data or new functionality? In this work, we study ways in which realistic user utterances can be generated synthetically, to help increase the linguistic and functional coverage of the system, without compromising the privacy of actual users. To this end, we propose a two-stage Differentially Private (DP) generation method which first generates latent semantic parses, and then generates utterances based on the parses. Our proposed approach improves MAUVE by 2.5X and parse tree function type overlap by 1.3X relative to current approaches for private synthetic data generation, improving both on fluency and semantic coverage. We further validate our approach on a realistic domain adaptation task of adding new functionality from private user data to a semantic parser, and show overall gains of 8.5% points in accuracy with the new feature.
翻译:面向任务的对话系统通常协助用户处理个人或机密事务。因此,此类系统的开发者通常被禁止观察实际使用情况。那么,他们如何知道系统在哪些方面存在不足,需要更多训练数据或新功能呢?在这项工作中,我们研究了如何合成生成真实的用户话语,以在不损害实际用户隐私的前提下,提高系统的语言和功能覆盖范围。为此,我们提出了一种两阶段差分隐私(DP)生成方法,该方法首先生成潜在语义解析,然后基于这些解析生成话语。与当前用于私有合成数据生成的方法相比,我们提出的方法在MAUVE指标上提升了2.5倍,在解析树功能类型重叠上提升了1.3倍,同时提高了流畅性和语义覆盖率。我们进一步在一个现实的领域自适应任务中验证了该方法——即从私有用户数据中向语义解析器添加新功能,结果显示新功能的准确率总体提升了8.5个百分点。