While recommender systems have become an integral component of the Web experience, their heavy reliance on user data raises privacy and security concerns. Substituting user data with synthetic data can address these concerns, but accurately replicating these real-world datasets has been a notoriously challenging problem. Recent advancements in generative AI have demonstrated the impressive capabilities of diffusion models in generating realistic data across various domains. In this work we introduce a Score-based Diffusion Recommendation Module (SDRM), which captures the intricate patterns of real-world datasets required for training highly accurate recommender systems. SDRM allows for the generation of synthetic data that can replace existing datasets to preserve user privacy, or augment existing datasets to address excessive data sparsity. Our method outperforms competing baselines such as generative adversarial networks, variational autoencoders, and recently proposed diffusion models in synthesizing various datasets to replace or augment the original data by an average improvement of 4.30% in Recall@$k$ and 4.65% in NDCG@$k$.
翻译:尽管推荐系统已成为网络体验的重要组成部分,但其对用户数据的高度依赖引发了隐私与安全方面的担忧。用合成数据替代用户数据可缓解这些问题,但精准复制这些真实世界数据集始终是一项极具挑战性的任务。生成式人工智能的最新进展展示了扩散模型在跨领域生成逼真数据方面的卓越能力。本研究提出了一种基于分数的扩散推荐模块(SDRM),该模块能够捕获训练高精度推荐系统所需的真实世界数据集中的复杂模式。SDRM可生成能替代现有数据集以保护用户隐私的合成数据,或增强现有数据集以应对数据过度稀疏问题。在合成多种数据集以替代或增强原始数据时,我们的方法在Recall@$k$和NDCG@$k$指标上分别平均提升4.30%和4.65%,性能优于生成对抗网络、变分自编码器及近期提出的扩散模型等竞争基线方法。