This paper introduces SynDiffix, a mechanism for generating statistically accurate, anonymous synthetic data for structured data. Recent open source and commercial systems use Generative Adversarial Networks or Transformed Auto Encoders to synthesize data, and achieve anonymity through overfitting-avoidance. By contrast, SynDiffix exploits traditional mechanisms of aggregation, noise addition, and suppression among others. Compared to CTGAN, ML models generated from SynDiffix are twice as accurate, marginal and column pairs data quality is one to two orders of magnitude more accurate, and execution time is two orders of magnitude faster. Compared to the best commercial product we measured (MostlyAI), ML model accuracy is comparable, marginal and pairs accuracy is 5 to 10 times better, and execution time is an order of magnitude faster. Similar to the other approaches, SynDiffix anonymization is very strong. This paper describes SynDiffix and compares its performance with other popular open source and commercial systems.
翻译:本文介绍了SynDiffix,一种用于生成统计精确、匿名的结构化数据合成机制。近期开源及商业系统多采用生成对抗网络(GAN)或变换自编码器(Transformed Auto Encoder)来合成数据,并通过避免过拟合实现匿名化。相比之下,SynDiffix利用了聚合、噪声添加与抑制等传统机制。与CTGAN相比,基于SynDiffix生成的机器学习模型精确度提高一倍,边际与列对数据质量提升一至两个数量级,执行速度加快两个数量级。与我们所测量的最佳商业产品(MostlyAI)相比,SynDiffix的机器学习模型精确度相当,边际与列对精确度提升5至10倍,执行速度加快一个数量级。与其他方法类似,SynDiffix的匿名化效果非常强。本文介绍了SynDiffix,并将其性能与其他主流开源及商业系统进行了比较。