Multi-armed bandit (MAB) processes constitute a foundational subclass of reinforcement learning problems and represent a central topic in statistical decision theory, but are limited to simultaneous adaptive allocation and sequential test, because of the absence of asymptotic theory under non-i.i.d sequence and sublinear information. To address this open challenge, we propose Urn Bandit (UNB) process to integrate the reinforcement mechanism of urn probabilistic models with MAB principles, ensuring almost sure convergence of resource allocation to optimal arms. We establish the joint functional central limit theorem (FCLT) for consistent estimators of expected rewards under non-i.i.d., non-sub-Gaussian and sublinear reward samples with pairwise correlations across arms. To overcome the limitations of existing methods that focus mainly on cumulative regret, we establish the asymptotic theory along with adaptive allocation that serves powerful sequential test, such as arms comparison, A/B testing, and policy valuation. Simulation studies and real data analysis demonstrate that UNB maintains statistical test performance of equal randomization (ER) design but obtain more average rewards like classical MAB processes.


翻译:多臂老虎机(MAB)过程是强化学习问题的一个基础子类,也是统计决策理论的核心课题,但由于在非独立同分布序列和次线性信息下缺乏渐近理论,其应用局限于同时进行自适应分配与序贯检验。为解决这一开放性挑战,我们提出瓮老虎机(UNB)过程,将瓮概率模型的强化机制与MAB原理相结合,确保资源分配几乎必然收敛至最优臂。针对臂间存在两两相关性、非独立同分布、非亚高斯且次线性的奖励样本,我们为期望奖励的一致估计量建立了联合泛函中心极限定理(FCLT)。为克服现有方法主要关注累积遗憾的局限,我们建立了结合自适应分配的渐近理论,该理论可为臂比较、A/B测试和策略评估等强大序贯检验提供支撑。仿真研究与实际数据分析表明,UNB在保持等随机化(ER)设计统计检验性能的同时,能获得类似经典MAB过程的更高平均奖励。

0
下载
关闭预览

相关内容

《分布式多智能体强化学习的编码》加州大学等
专知会员服务
57+阅读 · 2022年11月2日
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
强化学习初探 - 从多臂老虎机问题说起
专知
10+阅读 · 2018年4月3日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
博士论文 | 用代码结构感知方法推进代码大模型
《决策模型比较研究》
专知会员服务
9+阅读 · 7月25日
《美军水下战与海床战概述及本地实施》
专知会员服务
6+阅读 · 7月25日
面向未来冲突推进陆军情报体制改革
专知会员服务
4+阅读 · 7月25日
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
3+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
相关VIP内容
《分布式多智能体强化学习的编码》加州大学等
专知会员服务
57+阅读 · 2022年11月2日
相关基金
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员