Users' interaction or preference data used in recommender systems carry the risk of unintentionally revealing users' private attributes (e.g., gender or race). This risk becomes particularly concerning when the training data contains user preferences that can be used to infer these attributes, especially if they align with common stereotypes. This major privacy issue allows malicious attackers or other third parties to infer users' protected attributes. Previous efforts to address this issue have added or removed parts of users' preferences prior to or during model training to improve privacy, which often leads to decreases in recommendation accuracy. In this work, we introduce SBO, a novel probabilistic obfuscation method for user preference data designed to improve the accuracy--privacy trade-off for such recommendation scenarios. We apply SBO to three state-of-the-art recommendation models (i.e., BPR, MultVAE, and LightGCN) and two popular datasets (i.e., MovieLens-1M and LFM-2B). Our experiments reveal that SBO outperforms comparable approaches with respect to the accuracy--privacy trade-off. Specifically, we can reduce the leakage of users' protected attributes while maintaining on-par recommendation accuracy.
翻译:推荐系统中使用的用户交互或偏好数据存在无意间泄露用户隐私属性(如性别或种族)的风险。当训练数据包含可用于推断这些属性的用户偏好时(尤其是当这些偏好与常见刻板印象相符时),此风险尤为令人担忧。这一重大隐私问题使得恶意攻击者或其他第三方能够推断用户的受保护属性。先前解决此问题的方法是在模型训练前或训练期间添加或移除部分用户偏好以提升隐私性,但这通常会导致推荐准确性的下降。在本研究中,我们提出SBO——一种新颖的用户偏好数据概率混淆方法,旨在为此类推荐场景优化准确性-隐私性的权衡关系。我们将SBO应用于三种前沿推荐模型(即BPR、MultVAE和LightGCN)及两个常用数据集(即MovieLens-1M和LFM-2B)。实验结果表明,SBO在准确性-隐私性权衡方面优于同类方法。具体而言,我们能够在保持同等推荐准确性的同时,有效减少用户受保护属性的信息泄露。