Much of the research in differential privacy has focused on offline applications with the assumption that all data is available at once. When these algorithms are applied in practice to streams where data is collected over time, this either violates the privacy guarantees or results in poor utility. We derive an algorithm for differentially private synthetic streaming data generation, especially curated towards spatial datasets. Furthermore, we provide a general framework for online selective counting among a collection of queries which forms a basis for many tasks such as query answering and synthetic data generation. The utility of our algorithm is verified on both real-world and simulated datasets.
翻译:差分隐私领域的研究大多聚焦于离线应用,假设数据可一次性全部获得。当这些算法实际应用于随时间收集数据的流式场景时,要么会违反隐私保证,要么导致效用低下。我们提出了一种面向流式差分隐私合成数据生成的算法,尤其针对空间数据集进行了优化。此外,我们构建了一个通用框架,用于在查询集合中实现在线选择性计数,该框架为查询应答和合成数据生成等多项任务奠定了基础。我们在真实数据集与模拟数据集上验证了该算法的效用。