We introduce a novel extension of the canonical multi-armed bandit problem that incorporates an additional strategic element: abstention. In this enhanced framework, the agent is not only tasked with selecting an arm at each time step, but also has the option to abstain from accepting the stochastic instantaneous reward before observing it. When opting for abstention, the agent either suffers a fixed regret or gains a guaranteed reward. Given this added layer of complexity, we ask whether we can develop efficient algorithms that are both asymptotically and minimax optimal. We answer this question affirmatively by designing and analyzing algorithms whose regrets meet their corresponding information-theoretic lower bounds. Our results offer valuable quantitative insights into the benefits of the abstention option, laying the groundwork for further exploration in other online decision-making problems with such an option. Numerical results further corroborate our theoretical findings.
翻译:我们提出经典多臂赌博机问题的一个新扩展,该扩展引入了一个额外的策略性元素:弃权。在此增强框架中,智能体在每个时间步不仅需要选择一个臂,还可以在选择接受随机即时奖励之前选择弃权。当选择弃权时,智能体要么遭受固定的遗憾,要么获得有保证的奖励。鉴于这一增加的复杂性层面,我们探讨是否能够开发出既渐近最优又极小化极大最优的高效算法。我们通过设计和分析其遗憾值达到相应信息论下界的算法,对这一问题的回答是肯定的。我们的结果为弃权选项的益处提供了有价值的定量见解,为在其他具有此类选项的在线决策问题中的进一步探索奠定了基础。数值结果进一步证实了我们的理论发现。