Synthetic data has become a prominent solution for preserving privacy while sharing data, but current empirical risk assessment frameworks fundamentally assume a sample-based context that fails to translate for the evaluation of synthetic population level datasets. This commentary explores the implications when synthesizing entire populations in order to do population-level data science, arguing that traditional metrics, such as Membership Inference Attacks (MIA) and Attribute Inference Attacks (AIA), require re-examination. First, MIA may be rendered irrelevant in contexts where population membership is public knowledge or not considered sensitive information. Second, the risk of singling out is heightened because the confidential data contain full population information. Additionally, the absence of an "out-of-sample" comparison group for attribute inference means we need to define other policies when defining acceptable inferences. Finally, we cannot rely on simply returning to subsampling prior to generating synthetic data if the use case is truly to enable population-level data science. This commentary highlights the necessity for considering context when generating and evaluating synthetic population data.
翻译:合成数据已成为在共享数据时保护隐私的重要解决方案,但当前经验风险评估框架从根本上假设了基于样本的上下文,这种假设无法适用于合成人口级数据集的评估。本文探讨了为开展人口级数据科学而合成整个人口时的影响,论证了传统指标(如成员推理攻击和属性推理攻击)需要重新审视。首先,当人口成员身份属于公共知识或不被视为敏感信息时,成员推理攻击可能失去相关性。其次,由于机密数据包含完整的人口信息,识别特定个体(singling out)的风险会升高。此外,属性推理中"样本外"比较组的缺失意味着我们需要在定义可接受的推理时制定其他策略。最后,如果用例真正目的是实现人口级数据科学,我们不能简单依赖在生成合成数据前仅进行子采样。本文强调在生成和评估合成人口数据时考虑上下文环境的必要性。