Reinforcement learning (RL) has gained traction for enhancing user long-term experiences in recommender systems by effectively exploring users' interests. However, modern recommender systems exhibit distinct user behavioral patterns among tens of millions of items, which increases the difficulty of exploration. For example, user behaviors with different activity levels require varying intensity of exploration, while previous studies often overlook this aspect and apply a uniform exploration strategy to all users, which ultimately hurts user experiences in the long run. To address these challenges, we propose User-Oriented Exploration Policy (UOEP), a novel approach facilitating fine-grained exploration among user groups. We first construct a distributional critic which allows policy optimization under varying quantile levels of cumulative reward feedbacks from users, representing user groups with varying activity levels. Guided by this critic, we devise a population of distinct actors aimed at effective and fine-grained exploration within its respective user group. To simultaneously enhance diversity and stability during the exploration process, we further introduce a population-level diversity regularization term and a supervision module. Experimental results on public recommendation datasets demonstrate that our approach outperforms all other baselines in terms of long-term performance, validating its user-oriented exploration effectiveness. Meanwhile, further analyses reveal our approach's benefits of improved performance for low-activity users as well as increased fairness among users.
翻译:强化学习(RL)通过有效探索用户兴趣,已在提升推荐系统长期用户体验方面获得广泛关注。然而,现代推荐系统在数千万项目间呈现出显著的用户行为模式差异,这一特性增加了探索难度。例如,不同活跃度的用户行为需要差异化的探索强度,而现有研究往往忽视这一维度,对所有用户采用统一探索策略,最终损害长期用户体验。为应对这些挑战,我们提出面向用户的探索策略(UOEP),这是一种支持用户群体细粒度探索的创新方法。我们首先构建分布型评论家网络,该网络能够基于用户累积奖励反馈的不同分位水平进行策略优化,从而表征不同活跃度的用户群体。在此评论家网络指导下,我们设计了一组差异化执行器,旨在对各自用户群体实现高效且细粒度的探索。为同时增强探索过程中的多样性与稳定性,我们进一步引入群体级多样性正则化项和监督模块。在公开推荐数据集上的实验结果表明,本方法在长期性能方面优于所有基线模型,验证了其面向用户探索的有效性。同时,深入分析揭示了本方法在提升低活跃用户性能及增强用户间公平性方面的优势。