In the bandits with knapsacks framework (BwK) the learner has $m$ resource-consumption (packing) constraints. We focus on the generalization of BwK in which the learner has a set of general long-term constraints. The goal of the learner is to maximize their cumulative reward, while at the same time achieving small cumulative constraints violations. In this scenario, there exist simple instances where conventional methods for BwK fail to yield sublinear violations of constraints. We show that it is possible to circumvent this issue by requiring the primal and dual algorithm to be weakly adaptive. Indeed, even in absence on any information on the Slater's parameter $\rho$ characterizing the problem, the interplay between weakly adaptive primal and dual regret minimizers yields a "self-bounding" property of dual variables. In particular, their norm remains suitably upper bounded across the entire time horizon even without explicit projection steps. By exploiting this property, we provide best-of-both-worlds guarantees for stochastic and adversarial inputs. In the first case, we show that the algorithm guarantees sublinear regret. In the latter case, we establish a tight competitive ratio of $\rho/(1+\rho)$. In both settings, constraints violations are guaranteed to be sublinear in time. Finally, this results allow us to obtain new result for the problem of contextual bandits with linear constraints, providing the first no-$\alpha$-regret guarantees for adversarial contexts.
翻译:在背包约束强盗问题(BwK)中,学习器面临$m$个资源消耗(打包)约束。本文聚焦于BwK的泛化形式,其中学习器需处理一组一般性长期约束。学习器的目标是最大化累积奖励,同时实现累积约束违反量较小。在此场景下,存在若干简单实例,传统BwK方法无法获得约束违反量的次线性界。我们证明,通过要求原算法和对偶算法具有弱自适应性可规避此问题。事实上,即使在完全未知问题特征参数Slater系数$\rho$的情况下,弱自适应原遗憾最小化器与对偶遗憾最小化器之间的相互作用会引发对偶变量的“自界”特性。具体而言,即便未设置显式投影步骤,其对偶变量的范数在整个时间范围内仍能保持适当上界。通过利用这一特性,我们为随机输入与对抗输入提供了“两全其美”的保证:针对前者,算法可保证次线性遗憾;针对后者,我们建立了紧致竞争比$\rho/(1+\rho)$。在两种设定下,约束违反量均随时间保证为次线性。最终,这些结果使我们能够在线性约束的上下文强盗问题中获得新结论,首次为对抗性上下文提供无$\alpha$遗憾保证。