We consider a variant of contextual bandits in which the algorithm consumes multiple resources subject to linear constraints on total consumption. This problem generalizes contextual bandits with knapsacks (CBwK), allowing for packing and covering constraints, as well as positive and negative resource consumption. We present a new algorithm that is simple, computationally efficient, and admits vanishing regret. It is statistically optimal for CBwK when an algorithm must stop once some constraint is violated. Our algorithm builds on LagrangeBwK (Immorlica et al., FOCS 2019) , a Lagrangian-based technique for CBwK, and SquareCB (Foster and Rakhlin, ICML 2020), a regression-based technique for contextual bandits. Our analysis leverages the inherent modularity of both techniques.
翻译:我们考虑一种上下文赌博机的变体,其中算法在满足总消耗的线性约束条件下消耗多种资源。该问题推广了带背包的上下文赌博机(CBwK),允许打包约束和覆盖约束,以及正负资源消耗。我们提出一种新算法,该算法简单、计算高效且具有消失的遗憾值。当算法必须在某个约束被违反时终止时,该算法对CBwK在统计上是最优的。我们的算法建立在LagrangeBwK(Immorlica等人,FOCS 2019)——一种基于拉格朗日的CBwK技术——和SquareCB(Foster和Rakhlin,ICML 2020)——一种基于回归的上下文赌博机技术——之上。我们的分析利用了这两种技术固有的模块化特性。