This paper proposes a method of creating synthetic data (SD) that will have two important advantages for the user compared to other methods currently available. The first is transparency; unlike other methods, the person in receipt of the SD will know which of the relationships between variables in the original data will be approximately maintained in the SD. The second is a guarantee that the SD is derived from information that has already been judged to be free of disclosure risk. This is achieved by first defining and calculating the margins where relationships between variables will be maintained in the SD. Each margin will then be subject to statistical disclosure control (SDC) to the standards defined by the data custodian, e.g. top-coding and bottom-coding, combination of small categories and/or modifying small counts. Further adjustment of the curated margins is advised by coarsening all counts in the table to multiples of the disclosure limit. These adjusted margins are used to create SD by the Iterative Proportional Fitting (IPF) algorithm. The practical steps involved in creating such SD are illustrated using data from the 1901 Census of Scotland.
翻译:摘要:本文提出一种合成数据生成方法,相较于现有其他方法,该方法能为用户提供两大重要优势。其一是透明性:与其它方法不同,合成数据的接收方能够明确知晓原始数据中哪些变量间关系将在合成数据中被近似保留。其二是保证合成数据来源于已经认定无披露风险的信息。这一目标的实现首先需定义并计算变量间关系将在合成数据中被保留的边际。随后,每个边际将根据数据保管者制定的标准实施统计披露控制,例如顶端编码与底端编码、合并小类及/或修正小计数。建议将所有表格中的计数粗化至披露限制的整数倍,以进一步调整经过处理的边际。利用这些调整后的边际,通过迭代比例拟合算法生成合成数据。本文以1901年苏格兰人口普查数据为例,演示了生成此类合成数据所涉及的具体实践步骤。