Existing private synthetic data generation algorithms are agnostic to downstream tasks. However, end users may have specific requirements that the synthetic data must satisfy. Failure to meet these requirements could significantly reduce the utility of the data for downstream use. We introduce a post-processing technique that improves the utility of the synthetic data with respect to measures selected by the end user, while preserving strong privacy guarantees and dataset quality. Our technique involves resampling from the synthetic data to filter out samples that do not meet the selected utility measures, using an efficient stochastic first-order algorithm to find optimal resampling weights. Through comprehensive numerical experiments, we demonstrate that our approach consistently improves the utility of synthetic data across multiple benchmark datasets and state-of-the-art synthetic data generation algorithms.
翻译:现有的私有合成数据生成算法对下游任务无感知。然而,最终用户可能对合成数据需满足特定需求,若未能满足这些需求,将显著降低数据在下游使用中的效用。我们提出一种后处理技术,可在保持强隐私保障与数据集质量的前提下,针对用户选定的效用指标提升合成数据的实用性。该技术通过重采样合成数据来滤除不满足选定效用指标的样本,并采用高效的随机一阶算法求解最优重采样权重。通过综合数值实验表明,该方法在多个基准数据集及前沿合成数据生成算法上,均能持续提升合成数据的效用。