Recommender system always suffers from various recommendation biases, seriously hindering its development. In this light, a series of debias methods have been proposed in the recommender system, especially for two most common biases, i.e., popularity bias and amplified subjective bias. However, exsisting debias methods usually concentrate on correcting a single bias. Such single-functionality debiases neglect the bias-coupling issue in which the recommended items are collectively attributed to multiple biases. Besides, previous work cannot tackle the lacking supervised signals brought by sparse data, yet which has become a commonplace in the recommender system. In this work, we introduce a disentangled debias variational auto-encoder framework(DB-VAE) to address the single-functionality issue as well as a counterfactual data enhancement method to mitigate the adverse effect due to the data sparsity. In specific, DB-VAE first extracts two types of extreme items only affected by a single bias based on the collier theory, which are respectively employed to learn the latent representation of corresponding biases, thereby realizing the bias decoupling. In this way, the exact unbiased user representation can be learned by these decoupled bias representations. Furthermore, the data generation module employs Pearl's framework to produce massive counterfactual data, making up the lacking supervised signals due to the sparse data. Extensive experiments on three real-world datasets demonstrate the effectiveness of our proposed model. Besides, the counterfactual data can further improve DB-VAE, especially on the dataset with low sparsity.
翻译:推荐系统始终遭受各种推荐偏差的困扰,严重阻碍其发展。为此,一系列去偏方法已被提出,尤其针对两种最常见的偏差:流行度偏差和放大主观偏差。然而,现有去偏方法通常仅专注于纠正单一偏差。这种单功能去偏忽视了偏差耦合问题——推荐项目往往由多种偏差共同导致。此外,先前工作无法应对稀疏数据造成的监督信号缺失问题,而该问题在推荐系统中已十分普遍。本文提出一种解耦去偏变分自编码器框架(DB-VAE)以解决单功能问题,并引入反事实数据增强方法以缓解数据稀疏性带来的负面影响。具体而言,DB-VAE首先基于科利尔理论提取仅受单一偏差影响的两类极端项目,分别用于学习对应偏差的隐式表征,从而实现偏差解耦。通过这种方式,可基于解耦的偏差表征学习到精确的无偏用户表征。此外,数据生成模块采用Pearl框架生成海量反事实数据,弥补稀疏数据导致的监督信号缺失。在三个真实数据集上的大量实验证明了所提模型的有效性。同时,反事实数据可进一步提升DB-VAE的性能,尤其在低稀疏度数据集上效果显著。