We study the problem of efficiently generating differentially private synthetic data that approximate the statistical properties of an underlying sensitive dataset. In recent years, there has been a growing line of work that approaches this problem using first-order optimization techniques. However, such techniques are restricted to optimizing differentiable objectives only, severely limiting the types of analyses that can be conducted. For example, first-order mechanisms have been primarily successful in approximating statistical queries only in the form of marginals for discrete data domains. In some cases, one can circumvent such issues by relaxing the task's objective to maintain differentiability. However, even when possible, these approaches impose a fundamental limitation in which modifications to the minimization problem become additional sources of error. Therefore, we propose Private-GSD, a private genetic algorithm based on zeroth-order optimization heuristics that do not require modifying the original objective. As a result, it avoids the aforementioned limitations of first-order optimization. We empirically evaluate Private-GSD against baseline algorithms on data derived from the American Community Survey across a variety of statistics--otherwise known as statistical queries--both for discrete and real-valued attributes. We show that Private-GSD outperforms the state-of-the-art methods on non-differential queries while matching accuracy in approximating differentiable ones.
翻译:我们研究如何高效生成满足差分隐私的合成数据,使其近似逼近底层敏感数据集的统计特性。近年来,已有大量研究采用一阶优化技术处理该问题。然而,这类技术仅适用于优化可微目标函数,严重制约了可分析的数据类型。例如,一阶机制仅在离散数据域的边际统计查询中取得显著成效。在某些情况下,研究者可通过放宽任务目标以保持可微性来规避此类问题,但即便可行,这些方法仍存在根本性缺陷——对优化目标的修改会引入额外误差。为此,我们提出Private-GSD算法,这是一种基于零阶优化启发式策略的私有遗传算法,无需修改原始目标函数,从而避免了一阶优化的上述局限性。我们在来源于美国社区调查数据的多种统计量(即统计查询)上对Private-GSD与基线算法进行了实证评估,涵盖离散属性与实值属性。实验结果表明,Private-GSD在非可微查询上的性能超越当前最优方法,同时在近似可微查询时保持同等精度。