We study the corrupted bandit problem, i.e. a stochastic multi-armed bandit problem with $k$ unknown reward distributions, which are heavy-tailed and corrupted by a history-independent adversary or Nature. To be specific, the reward obtained by playing an arm comes from corresponding heavy-tailed reward distribution with probability $1-\varepsilon \in (0.5,1]$ and an arbitrary corruption distribution of unbounded support with probability $\varepsilon \in [0,0.5)$. First, we provide $\textit{a problem-dependent lower bound on the regret}$ of any corrupted bandit algorithm. The lower bounds indicate that the corrupted bandit problem is harder than the classical stochastic bandit problem with sub-Gaussian or heavy-tail rewards. Following that, we propose a novel UCB-type algorithm for corrupted bandits, namely HubUCB, that builds on Huber's estimator for robust mean estimation. Leveraging a novel concentration inequality of Huber's estimator, we prove that HubUCB achieves a near-optimal regret upper bound. Since computing Huber's estimator has quadratic complexity, we further introduce a sequential version of Huber's estimator that exhibits linear complexity. We leverage this sequential estimator to design SeqHubUCB that enjoys similar regret guarantees while reducing the computational burden. Finally, we experimentally illustrate the efficiency of HubUCB and SeqHubUCB in solving corrupted bandits for different reward distributions and different levels of corruptions.


翻译:我们研究自然败坏的赌博机问题,即具有$k$个未知奖励分布的随机多臂赌博机问题,这些分布具有重尾特性,且被历史无关的对手或自然所败坏。具体而言,通过拉动臂获得的奖励以概率$1-\varepsilon \in (0.5,1]$来自对应重尾奖励分布,并以概率$\varepsilon \in [0,0.5)$来自具有无界支撑的任意败坏分布。首先,我们给出任何败坏赌博机算法的$\textit{问题依赖的遗憾下界}$。下界表明,与具有次高斯或重尾奖励的经典随机赌博机问题相比,自然败坏的赌博机问题更难解决。随后,针对自然败坏的赌博机,我们提出一种新颖的基于UCB的算法——HubUCB,它构建在用于鲁棒均值估计的Huber估计器之上。利用Huber估计器的新颖浓度不等式,我们证明HubUCB实现了近乎最优的遗憾上界。由于计算Huber估计器具有二次复杂度,我们进一步引入一种具有线性复杂度的序贯版本Huber估计器。利用该序贯估计器,我们设计出SeqHubUCB算法,其在减少计算负担的同时享有类似的遗憾保证。最后,我们通过实验展示了HubUCB和SeqHubUCB在解决不同奖励分布和不同败坏程度下自然败坏的赌博机问题中的效率。

0
下载
关闭预览

相关内容

【ICML2023】序列反事实风险最小化
专知会员服务
21+阅读 · 2023年5月1日
《分布式多智能体强化学习的编码》加州大学等
专知会员服务
57+阅读 · 2022年11月2日
【ICML2022】鲁棒强化学习的策略梯度法
专知会员服务
38+阅读 · 2022年5月21日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
21+阅读 · 2021年10月24日
专知会员服务
41+阅读 · 2021年2月12日
Fariz Darari简明《博弈论Game Theory》介绍,35页ppt
专知会员服务
113+阅读 · 2020年5月15日
专知会员服务
63+阅读 · 2020年3月4日
一些关于随机矩阵的算法
PaperWeekly
1+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证
深度强化学习实验室
20+阅读 · 2020年8月11日
RL解决'LunarLander-v2' (SOTA)
CreateAMind
62+阅读 · 2019年9月27日
【重磅】61篇NIPS2019深度强化学习论文及部分解读
机器学习算法与Python学习
10+阅读 · 2019年9月14日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Arxiv
0+阅读 · 2023年5月12日
Arxiv
0+阅读 · 2023年5月12日
Arxiv
0+阅读 · 2023年5月10日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
2+阅读 · 今天7:07
《无人机空中监控:通信实验洞察》
专知会员服务
1+阅读 · 今天7:05
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
5+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
5+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
相关VIP内容
【ICML2023】序列反事实风险最小化
专知会员服务
21+阅读 · 2023年5月1日
《分布式多智能体强化学习的编码》加州大学等
专知会员服务
57+阅读 · 2022年11月2日
【ICML2022】鲁棒强化学习的策略梯度法
专知会员服务
38+阅读 · 2022年5月21日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
21+阅读 · 2021年10月24日
专知会员服务
41+阅读 · 2021年2月12日
Fariz Darari简明《博弈论Game Theory》介绍,35页ppt
专知会员服务
113+阅读 · 2020年5月15日
专知会员服务
63+阅读 · 2020年3月4日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员