Task-oriented dialogue systems often assist users with personal or confidential matters. For this reason, the developers of such a system are generally prohibited from observing actual usage. So how can they know where the system is failing and needs more training data or new functionality? In this work, we study ways in which realistic user utterances can be generated synthetically, to help increase the linguistic and functional coverage of the system, without compromising the privacy of actual users. To this end, we propose a two-stage Differentially Private (DP) generation method which first generates latent semantic parses, and then generates utterances based on the parses. Our proposed approach improves MAUVE by 2.5$\times$ and parse tree function type overlap by 1.3$\times$ relative to current approaches for private synthetic data generation, improving both on fluency and semantic coverage. We further validate our approach on a realistic domain adaptation task of adding new functionality from private user data to a semantic parser, and show overall gains of 8.5% points in accuracy with the new feature.
翻译:任务导向型对话系统常协助用户处理个人或机密事务。因此,此类系统的开发者通常被禁止观察实际使用情况。那么,他们如何知晓系统在何处出现故障、需要更多训练数据或新功能?在本研究中,我们探索通过合成方式生成逼真的用户语句的方法,以在不损害实际用户隐私的前提下,提升系统的语言与功能覆盖范围。为此,我们提出一种两阶段差分隐私生成方法:首先生成潜在语义解析,再基于这些解析生成语句。与当前用于私有合成数据生成的方法相比,我们所提出的方法将MAUVE指标提升了2.5倍,解析树函数类型重叠提升了1.3倍,并在流畅性和语义覆盖率两方面均有所改善。我们进一步在一个现实领域自适应任务中验证了该方法——即从私有用户数据中为语义解析器添加新功能——结果显示,新功能的准确率整体提升了8.5个百分点。