Electronic Health Records (EHRs) contain sensitive patient information, which presents privacy concerns when sharing such data. Synthetic data generation is a promising solution to mitigate these risks, often relying on deep generative models such as Generative Adversarial Networks (GANs). However, recent studies have shown that diffusion models offer several advantages over GANs, such as generation of more realistic synthetic data and stable training in generating data modalities, including image, text, and sound. In this work, we investigate the potential of diffusion models for generating realistic mixed-type tabular EHRs, comparing TabDDPM model with existing methods on four datasets in terms of data quality, utility, privacy, and augmentation. Our experiments demonstrate that TabDDPM outperforms the state-of-the-art models across all evaluation metrics, except for privacy, which confirms the trade-off between privacy and utility.
翻译:电子健康记录(EHRs)包含敏感的患者信息,在共享此类数据时会引发隐私问题。合成数据生成是缓解这些风险的一种有前景的解决方案,通常依赖于深度生成模型,如生成对抗网络(GANs)。然而,近期研究表明,扩散模型相比GANs具有多项优势,例如能生成更真实的合成数据,并在图像、文本和声音等多种数据模态生成过程中实现稳定训练。在本研究中,我们探讨了扩散模型在生成真实混合类型表格型EHRs方面的潜力,从数据质量、实用性、隐私保护和数据增强四个维度,将TabDDPM模型与现有方法在四个数据集上进行对比。实验结果表明,除隐私保护指标外,TabDDPM在所有评估指标上均优于当前最优模型,这验证了隐私性与实用性之间的权衡关系。