Synthetic data is emerging as a promising way to harness the value of data, while reducing privacy risks. The potential of synthetic data is not limited to privacy-friendly data release, but also includes complementing real data in use-cases such as training machine learning algorithms that are more fair and robust to distribution shifts etc. There is a lot of interest in algorithmic advances in synthetic data generation for providing better privacy and statistical guarantees and for its better utilisation in machine learning pipelines. However, for responsible and trustworthy synthetic data generation, it is not sufficient to focus only on these algorithmic aspects and instead, a holistic view of the synthetic data generation pipeline must be considered. We build a novel system that allows the contributors of real data to autonomously participate in differentially private synthetic data generation without relying on a trusted centre. Our modular, general and scalable solution is based on three building blocks namely: Solid (Social Linked Data), MPC (Secure Multi-Party Computation) and Trusted Execution Environments (TEEs). Solid is a specification that lets people store their data securely in decentralised data stores called Pods and control access to their data. MPC refers to the set of cryptographic methods for different parties to jointly compute a function over their inputs while keeping those inputs private. TEEs such as Intel SGX rely on hardware based features for confidentiality and integrity of code and data. We show how these three technologies can be effectively used to address various challenges in responsible and trustworthy synthetic data generation by ensuring: 1) contributor autonomy, 2) decentralisation, 3) privacy and 4) scalability. We support our claims with rigorous empirical results on simulated and real datasets and different synthetic data generation algorithms.
翻译:综合数据生成正成为一种有前景的方式,既能利用数据的价值,又能降低隐私风险。综合数据的潜力不仅局限于隐私友好的数据发布,还包括在训练更公平、对分布变化更具鲁棒性的机器学习算法等用例中补充真实数据。为在综合数据生成中提供更好的隐私和统计保证,并使其在机器学习流程中得到更优利用,相关算法研究备受关注。然而,要实现负责任且可信赖的综合数据生成,仅关注这些算法方面是不够的,必须从整体视角审视综合数据生成流程。我们构建了一套新颖系统,使真实数据的贡献者能够自主参与差分隐私综合数据生成,而无需依赖可信中心。我们模块化、通用且可扩展的解决方案基于三大组件:Solid(社交关联数据)、MPC(安全多方计算)和可信执行环境(TEE)。Solid是一种规范,允许人们将数据安全地存储在名为Pods的去中心化数据存储中,并控制对其数据的访问。MPC指一组密码学方法,使不同参与方能在保持输入私密的同时共同计算函数。TEE(如Intel SGX)依赖基于硬件的功能确保代码和数据的机密性与完整性。我们展示了如何有效利用这三种技术,通过确保以下特性来应对负责任且可信赖的综合数据生成中的各种挑战:1)贡献者自主性,2)去中心化,3)隐私性,4)可扩展性。我们通过在模拟和真实数据集上使用不同综合数据生成算法进行的严格实证结果支持了这些论断。