Recommender systems play a crucial role in internet economies by connecting users with relevant products. However, designing effective recommender systems faces the key challenges: the exploration-exploitation tradeoff in securing incentive to explore new products against user's self-interested preferences. While prior work addresses Bayesian Incentive Compatibility (BIC) in fixed-design linear bandits (Sellke & Slivkins, 2023), we tackle the challenge of stochastic user covariates sampled online. Unlike standard black-box reductions (Mansour et al., 2020), our two-stage framework exploits the linear reward structure to achieve sublinear regret while satisfying incentive constraints. To address it, we propose a two-stage algorithm that integrates incentivized exploration with any efficient plug-in offline learning algorithms. In the first stage, it explores products while maintaining incentive compatibility to gather optimal samples. The second stage employs inverse proportional gap sampling strategy (IPGS) integrated with any efficient learning methods to secure sublinear regret. Theoretically, we prove that algorithm RCB achieves $O(\sqrt{KdT})$ regret and simultaneously satisfies incentive constraints, and discovers the tradeoff between incentive budget and regret, validating in experiments. We demonstrate RCB's strong incentive gain, sublinear regret, and robustness through a real application on personalized warfarin dosing and simulations.
翻译:推荐系统在互联网经济中发挥着连接用户与相关产品的关键作用。然而,设计有效的推荐系统面临着核心挑战:在应对用户自利偏好的同时,如何通过探索-利用权衡来确保探索新产品的激励。现有研究基于固定设计线性bandit实现了贝叶斯激励相容性(Sellke & Slivkins, 2023),而本文则应对在线随机采样用户协变量的挑战。与标准黑盒归约方法(Mansour等人,2020)不同,我们的两阶段框架通过利用线性奖励结构,在满足激励约束的同时实现次线性遗憾。为此,我们提出一种两阶段算法,将激励探索与任何高效即插式离线学习算法相融合。第一阶段在保持激励相容性的同时探索产品以收集最优样本;第二阶段采用逆比例间隙采样策略(IPGS)与高效学习方法集成,确保次线性遗憾。理论上,我们证明RCB算法可实现$O(\sqrt{KdT})$遗憾值并同时满足激励约束,揭示了激励预算与遗憾值之间的权衡关系,并通过实验验证。我们在个性化华法林剂量推荐的真实应用及仿真实验中,展示了RCB算法强大的激励增益、次线性遗憾值及鲁棒性。