Synthetic data holds substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and complexity for non-experts. Furthermore, common evaluation strategies for synthetic data often fail to directly reflect statistical utility and measure privacy risks sufficiently. Against this background, a critical underexplored question is whether synthetic data can reliably reproduce key findings from epidemiological research while preserving privacy. We propose adversarial random forests (ARF) as an efficient and convenient method for synthesizing tabular epidemiological data. To evaluate its performance, we replicated statistical analyses from six epidemiological publications covering blood pressure, anthropometry, myocardial infarction, accelerometry, loneliness, and diabetes, from the German National Cohort (NAKO Gesundheitsstudie), the Bremen STEMI Registry U45 Study, and the Guelph Family Health Study. We further assessed how dataset dimensionality and variable complexity affect the quality of synthetic data, and contextualized ARF's performance by comparison with commonly used tabular data synthesizers in terms of utility, privacy, generalisation, and runtime. Across all replicated studies, results on ARF-generated synthetic data consistently aligned with original findings. Even for datasets with relatively low sample size-to-dimensionality ratios, replication outcomes closely matched the original results across descriptive and inferential analyses. Reduced dimensionality and variable complexity further enhanced synthesis quality. ARF demonstrated favourable performance regarding utility, privacy preservation, and generalisation relative to other synthesizers and superior computational efficiency.
翻译:合成数据在应对流行病学中因数据访问受限和隐私问题产生的实际挑战方面具有巨大潜力。然而,当前许多方法存在质量有限、计算需求高、对非专业人员复杂等问题。此外,常见的合成数据评估策略往往不能直接反映统计效用,也难以充分衡量隐私风险。在此背景下,一个关键但尚未充分探索的问题是:合成数据能否在保护隐私的同时,可靠地重现流行病学研究的关键发现?我们提出使用对抗随机森林(ARF)作为一种高效且便捷的表格型流行病学数据合成方法。为评估其性能,我们复制了来自德国国家队列(NAKO Gesundheitsstudie)、不来梅STEMI登记U45研究以及圭尔夫家庭健康研究中六篇流行病学出版物(涵盖血压、人体测量学、心肌梗死、加速度测量、孤独感和糖尿病)的统计分析。我们进一步评估了数据集维度和变量复杂性如何影响合成数据的质量,并通过与常用表格型数据合成器在效用、隐私、泛化能力和运行时间方面的比较,对ARF的性能进行了情境化分析。在所有复制研究中,基于ARF生成合成数据所得出的结果均与原始发现保持一致。即使在样本量与维度之比较低的数据集中,描述性和推断性分析的复制结果也与原始结果高度吻合。降低数据维度和变量复杂性进一步提升了合成质量。与其他合成器相比,ARF在效用、隐私保护和泛化能力方面表现出色,并具有优越的计算效率。