Simple regret minimization is a critical problem in learning optimal treatment assignment policies across various domains, including healthcare and e-commerce. However, it remains understudied in the contextual bandit setting. We propose a new family of computationally efficient bandit algorithms for the stochastic contextual bandit settings, with the flexibility to be adapted for cumulative regret minimization (with near-optimal minimax guarantees) and simple regret minimization (with SOTA guarantees). Furthermore, our algorithms adapt to model misspecification and extend to the continuous arm settings. These advantages come from constructing and relying on "conformal arm sets" (CASs), which provide a set of arms at every context that encompass the context-specific optimal arm with some probability across the context distribution. Our positive results on simple and cumulative regret guarantees are contrasted by a negative result, which shows that an algorithm can't achieve instance-dependent simple regret guarantees while simultaneously achieving minimax optimal cumulative regret guarantees.
翻译:简单遗憾最小化是学习最优治疗分配策略的关键问题,广泛应用于医疗保健和电子商务等领域。然而,该问题在上下文赌博机场景中的研究仍显不足。我们针对随机上下文赌博机场景提出一类新的计算高效型赌博机算法,其灵活性可同时适用于累积遗憾最小化(具备接近最优的极小化极大保证)和简单遗憾最小化(具备最新最优保证)。此外,我们的算法能适应模型误设问题,并扩展至连续臂设置。这些优势源于构造并依赖"一致性臂集"——该集合能在每个上下文中提供一组臂,使得该集合以一定概率(基于上下文分布)包含对应上下文的最优臂。我们在简单与累积遗憾保证方面的积极结果与一个消极结果形成对比:该消极结果表明,不存在一种算法能同时实现实例相关的简单遗憾保证和极小化极大最优的累积遗憾保证。