We propose a novel approach for developing privacy-preserving large-scale recommender systems using differentially private (DP) large language models (LLMs) which overcomes certain challenges and limitations in DP training these complex systems. Our method is particularly well suited for the emerging area of LLM-based recommender systems, but can be readily employed for any recommender systems that process representations of natural language inputs. Our approach involves using DP training methods to fine-tune a publicly pre-trained LLM on a query generation task. The resulting model can generate private synthetic queries representative of the original queries which can be freely shared for any downstream non-private recommendation training procedures without incurring any additional privacy cost. We evaluate our method on its ability to securely train effective deep retrieval models, and we observe significant improvements in their retrieval quality without compromising query-level privacy guarantees compared to methods where the retrieval models are directly DP trained.
翻译:本文提出了一种利用差分隐私(DP)大语言模型(LLM)开发隐私保护大规模推荐系统的新方法,该方法克服了在DP训练这些复杂系统时所面临的部分挑战与局限性。我们的方法特别适用于基于LLM的推荐系统这一新兴领域,但也可轻松应用于任何处理自然语言输入表征的推荐系统。该方法采用DP训练技术,对公开预训练的LLM进行查询生成任务的微调。由此生成的模型能够产生代表原始查询的私有合成查询,这些查询可自由共享用于任何下游非隐私保护的推荐训练流程,且不会产生额外的隐私成本。我们在安全训练高效深度检索模型的能力上评估了该方法,与直接对检索模型进行DP训练的方法相比,本方法在检索质量上实现了显著提升,同时未损害查询级别的隐私保障。