Recommender systems play a crucial role in internet economies by connecting users with relevant products. However, designing effective recommender systems faces the key challenges: the exploration-exploitation tradeoff in securing incentive to explore new products against user's self-interested preferences. While prior work addresses Bayesian Incentive Compatibility (BIC) in fixed-design linear bandits (Sellke & Slivkins, 2023), we tackle the challenge of stochastic user covariates sampled online. Unlike standard black-box reductions (Mansour et al., 2020), our two-stage framework exploits the linear reward structure to achieve sublinear regret while satisfying incentive constraints. To address it, we propose a two-stage algorithm that integrates incentivized exploration with any efficient plug-in offline learning algorithms. In the first stage, it explores products while maintaining incentive compatibility to gather optimal samples. The second stage employs inverse proportional gap sampling strategy (IPGS) integrated with any efficient learning methods to secure sublinear regret. Theoretically, we prove that algorithm RCB achieves $O(\sqrt{KdT})$ regret and simultaneously satisfies incentive constraints, and discovers the tradeoff between incentive budget and regret, validating in experiments. We demonstrate RCB's strong incentive gain, sublinear regret, and robustness through a real application on personalized warfarin dosing and simulations.


翻译:推荐系统在互联网经济中发挥着连接用户与相关产品的关键作用。然而,设计有效的推荐系统面临着核心挑战:在应对用户自利偏好的同时,如何通过探索-利用权衡来确保探索新产品的激励。现有研究基于固定设计线性bandit实现了贝叶斯激励相容性(Sellke & Slivkins, 2023),而本文则应对在线随机采样用户协变量的挑战。与标准黑盒归约方法(Mansour等人,2020)不同,我们的两阶段框架通过利用线性奖励结构,在满足激励约束的同时实现次线性遗憾。为此,我们提出一种两阶段算法,将激励探索与任何高效即插式离线学习算法相融合。第一阶段在保持激励相容性的同时探索产品以收集最优样本;第二阶段采用逆比例间隙采样策略(IPGS)与高效学习方法集成,确保次线性遗憾。理论上,我们证明RCB算法可实现$O(\sqrt{KdT})$遗憾值并同时满足激励约束,揭示了激励预算与遗憾值之间的权衡关系,并通过实验验证。我们在个性化华法林剂量推荐的真实应用及仿真实验中,展示了RCB算法强大的激励增益、次线性遗憾值及鲁棒性。

0
下载
关闭预览

相关内容

设计是对现有状的一种重新认识和打破重组的过程,设计让一切变得更美。
基于因果推断的推荐系统去偏研究
专知会员服务
21+阅读 · 2024年11月10日
自监督学习推荐系统综述
专知会员服务
37+阅读 · 2024年4月6日
对话推荐算法研究综述
专知会员服务
50+阅读 · 2022年2月18日
个性化推荐系统技术进展
专知会员服务
66+阅读 · 2020年8月15日
基于知识图谱的推荐系统研究综述
专知会员服务
334+阅读 · 2020年8月10日
推荐系统(一):推荐系统基础
菜鸟的机器学习
25+阅读 · 2019年9月2日
新书推荐《推荐系统进展:方法与技术》
LibRec智能推荐
13+阅读 · 2019年3月18日
详解 | 推荐系统的工程实现
AI100
42+阅读 · 2019年3月15日
推荐系统
炼数成金订阅号
28+阅读 · 2019年1月17日
推荐系统概述
Linux爱好者
20+阅读 · 2018年9月6日
深度学习在推荐系统中的应用综述(最全)
七月在线实验室
17+阅读 · 2018年5月5日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月12日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
8+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
8+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
11+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
12+阅读 · 8月8日
相关资讯
推荐系统(一):推荐系统基础
菜鸟的机器学习
25+阅读 · 2019年9月2日
新书推荐《推荐系统进展:方法与技术》
LibRec智能推荐
13+阅读 · 2019年3月18日
详解 | 推荐系统的工程实现
AI100
42+阅读 · 2019年3月15日
推荐系统
炼数成金订阅号
28+阅读 · 2019年1月17日
推荐系统概述
Linux爱好者
20+阅读 · 2018年9月6日
深度学习在推荐系统中的应用综述(最全)
七月在线实验室
17+阅读 · 2018年5月5日
相关基金
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员