We introduce Coarse Q-learning (CQL), a reinforcement-learning model for bandit problems with stochastically varying menus. Alternatives are exogenously partitioned into similarity classes, and feedback from sampled alternatives is pooled within classes into class-level valuations. Choices follow multinomial logit over class valuations, and valuations update toward realized payoffs as in Q-learning. Using stochastic approximation, we derive the mean-field dynamics and characterize the steady states as smooth analogues of Valuation Equilibria. The model yields novel long-run phenomena in the high payoff-sensitivity limit: depending on the environment, CQL may exhibit multiple stable strict equilibria, a unique globally stable mixed equilibrium with indifference across classes, or no stable equilibrium at all, with valuations and choice probabilities converging instead to a stable limit cycle. These outcomes are driven by coarse aggregation and do not arise in the standard alternative-level benchmark.


翻译:我们提出粗Q学习(CQL),一种针对菜单随机变化的赌博机问题的强化学习模型。备选方案被外生地划分为相似性类别,采样备选方案的反馈在类别内聚合为类别级估值。选择遵循基于类别估值的多项Logit模型,估值像Q学习一样向实现收益更新。利用随机逼近方法,我们推导出平均场动力学,并将稳态刻画为估值均衡的光滑类比。模型在高收益敏感性极限下产生了新颖的长期现象:根据环境不同,CQL可能呈现多重稳定严格均衡、所有类别间漠然的唯一全局稳定混合均衡,或根本不存在稳定均衡——此时估值和选择概率收敛于稳定极限环。这些结果由粗粒化聚合驱动,在标准备选方案级基准中不会出现。

0
下载
关闭预览

相关内容

论学习、公平性与复杂度
专知会员服务
12+阅读 · 2月28日
【ACL2024】语言模型对齐的不确定性感知学习
专知会员服务
26+阅读 · 2024年6月10日
专知会员服务
17+阅读 · 2020年12月4日
【MIT】反偏差对比学习,Debiased Contrastive Learning
专知会员服务
92+阅读 · 2020年7月4日
强化学习开篇:Q-Learning原理详解
AINLP
37+阅读 · 2020年7月28日
浅谈主动学习(Active Learning)
凡人机器学习
32+阅读 · 2020年6月18日
元强化学习迎来一盆冷水:不比元Q学习好多少
AI科技评论
12+阅读 · 2020年2月27日
强化学习扫盲贴:从Q-learning到DQN
夕小瑶的卖萌屋
52+阅读 · 2019年10月13日
无监督元学习表示学习
CreateAMind
27+阅读 · 2019年1月4日
入门 | 通过 Q-learning 深入理解强化学习
机器之心
12+阅读 · 2018年4月17日
一个强化学习 Q-learning 算法的简明教程
数据挖掘入门与实战
10+阅读 · 2018年3月18日
入门 | 从Q学习到DDPG,一文简述多种强化学习算法
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
VIP会员
最新内容
《基于强化学习的自动化红队测试》
专知会员服务
3+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
3+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
7+阅读 · 7月20日
相关资讯
强化学习开篇:Q-Learning原理详解
AINLP
37+阅读 · 2020年7月28日
浅谈主动学习(Active Learning)
凡人机器学习
32+阅读 · 2020年6月18日
元强化学习迎来一盆冷水:不比元Q学习好多少
AI科技评论
12+阅读 · 2020年2月27日
强化学习扫盲贴:从Q-learning到DQN
夕小瑶的卖萌屋
52+阅读 · 2019年10月13日
无监督元学习表示学习
CreateAMind
27+阅读 · 2019年1月4日
入门 | 通过 Q-learning 深入理解强化学习
机器之心
12+阅读 · 2018年4月17日
一个强化学习 Q-learning 算法的简明教程
数据挖掘入门与实战
10+阅读 · 2018年3月18日
入门 | 从Q学习到DDPG,一文简述多种强化学习算法
相关基金
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员