Motivated by privacy concerns in long-term longitudinal studies in medical and social science research, we study the problem of continually releasing differentially private synthetic data. We introduce a model where, in every time step, each individual reports a new data element, and the goal of the synthesizer is to incrementally update a synthetic dataset to capture a rich class of statistical properties. We give continual synthetic data generation algorithms that preserve two basic types of queries: fixed time window queries and cumulative time queries. We show nearly tight upper bounds on the error rates of these algorithms and demonstrate their empirical performance on realistically sized datasets from the U.S. Census Bureau's Survey of Income and Program Participation.
翻译:受医学和社会科学研究中长期纵向研究中隐私问题的驱动,我们研究持续发布差分隐私合成数据的问题。我们引入一个模型,其中在每个时间步,每个个体报告一个新的数据元素,而合成器的目标是逐步更新合成数据集以捕获丰富的统计特性类别。我们给出了持续合成数据生成算法,这些算法保留了两类基本查询:固定时间窗口查询和累积时间查询。我们展示了这些算法误差率几乎严格的上界,并在美国人口调查局收入与项目参与调查的真实规模数据集上验证了其实验性能。