Synthetic data holds substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and complexity for non-experts. Furthermore, common evaluation strategies for synthetic data often fail to directly reflect statistical utility and measure privacy risks sufficiently. Against this background, a critical underexplored question is whether synthetic data can reliably reproduce key findings from epidemiological research while preserving privacy. We propose adversarial random forests (ARF) as an efficient and convenient method for synthesizing tabular epidemiological data. To evaluate its performance, we replicated statistical analyses from six epidemiological publications covering blood pressure, anthropometry, myocardial infarction, accelerometry, loneliness, and diabetes, from the German National Cohort (NAKO Gesundheitsstudie), the Bremen STEMI Registry U45 Study, and the Guelph Family Health Study. We further assessed how dataset dimensionality and variable complexity affect the quality of synthetic data, and contextualized ARF's performance by comparison with commonly used tabular data synthesizers in terms of utility, privacy, generalisation, and runtime. Across all replicated studies, results on ARF-generated synthetic data consistently aligned with original findings. Even for datasets with relatively low sample size-to-dimensionality ratios, replication outcomes closely matched the original results across descriptive and inferential analyses. Reduced dimensionality and variable complexity further enhanced synthesis quality. ARF demonstrated favourable performance regarding utility, privacy preservation, and generalisation relative to other synthesizers and superior computational efficiency.


翻译:合成数据在应对流行病学中因数据访问受限和隐私问题产生的实际挑战方面具有巨大潜力。然而,当前许多方法存在质量有限、计算需求高、对非专业人员复杂等问题。此外,常见的合成数据评估策略往往不能直接反映统计效用,也难以充分衡量隐私风险。在此背景下,一个关键但尚未充分探索的问题是:合成数据能否在保护隐私的同时,可靠地重现流行病学研究的关键发现?我们提出使用对抗随机森林(ARF)作为一种高效且便捷的表格型流行病学数据合成方法。为评估其性能,我们复制了来自德国国家队列(NAKO Gesundheitsstudie)、不来梅STEMI登记U45研究以及圭尔夫家庭健康研究中六篇流行病学出版物(涵盖血压、人体测量学、心肌梗死、加速度测量、孤独感和糖尿病)的统计分析。我们进一步评估了数据集维度和变量复杂性如何影响合成数据的质量,并通过与常用表格型数据合成器在效用、隐私、泛化能力和运行时间方面的比较,对ARF的性能进行了情境化分析。在所有复制研究中,基于ARF生成合成数据所得出的结果均与原始发现保持一致。即使在样本量与维度之比较低的数据集中,描述性和推断性分析的复制结果也与原始结果高度吻合。降低数据维度和变量复杂性进一步提升了合成质量。与其他合成器相比,ARF在效用、隐私保护和泛化能力方面表现出色,并具有优越的计算效率。

0
下载
关闭预览

相关内容

《利用合成数据生成加强军事决策支持》
专知会员服务
43+阅读 · 2024年12月30日
【MIT博士论文】合成数据的视觉表示学习
专知会员服务
27+阅读 · 2024年8月25日
谷歌最新《大语言模型合成数据的最佳实践和经验教训》
流行病数据可视分析综述
专知会员服务
27+阅读 · 2022年3月21日
图像修复研究进展综述
专知
20+阅读 · 2021年3月9日
基于深度学习的数据融合方法研究综述
专知
37+阅读 · 2020年12月10日
用深度学习揭示数据的因果关系
专知
28+阅读 · 2019年5月18日
视频生成的前沿论文,看我们推荐的7篇就够了
人工智能前沿讲习班
34+阅读 · 2018年12月30日
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
Arxiv
0+阅读 · 4月10日
VIP会员
相关主题
最新内容
面向2027年及未来的海军情报改革
专知会员服务
3+阅读 · 8月5日
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
4+阅读 · 8月5日
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 8月5日
相关基金
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员