The idea to generate synthetic data as a tool for broadening access to sensitive microdata has been proposed for the first time three decades ago. While first applications of the idea emerged around the turn of the century, the approach really gained momentum over the last ten years, stimulated at least in parts by some recent developments in computer science. We consider the upcoming 30th jubilee of Rubin's seminal paper on synthetic data (Rubin, 1993) as an opportunity to look back at the historical developments, but also to offer a review of the diverse approaches and methodological underpinnings proposed over the years. We will also discuss the various strategies that have been suggested to measure the utility and remaining risk of disclosure of the generated data.
翻译:首次提出生成合成数据作为拓宽敏感微观数据访问工具这一构想,距今已三十年。虽然该构想的首批应用出现在世纪之交,但近十年来,在计算机科学最新发展的部分推动下,该方法才真正获得发展动力。我们以鲁宾(Rubin, 1993)关于合成数据开创性论文即将迎来三十周年纪念为契机,回顾这一领域的历史发展脉络,同时综述多年来提出的各类方法及其方法论基础。此外,本文还将探讨用于评估生成数据效用及剩余披露风险的各类策略。