We consider a combinatorial multi-armed bandit problem for maximum value reward function under maximum value and index feedback. This is a new feedback structure that lies in between commonly studied semi-bandit and full-bandit feedback structures. We propose an algorithm and provide a regret bound for problem instances with stochastic arm outcomes according to arbitrary distributions with finite supports. The regret analysis rests on considering an extended set of arms, associated with values and probabilities of arm outcomes, and applying a smoothness condition. Our algorithm achieves a $O((k/\Delta)\log(T))$ distribution-dependent and a $\tilde{O}(\sqrt{T})$ distribution-independent regret where $k$ is the number of arms selected in each round, $\Delta$ is a distribution-dependent reward gap and $T$ is the horizon time. Perhaps surprisingly, the regret bound is comparable to previously-known bound under more informative semi-bandit feedback. We demonstrate the effectiveness of our algorithm through experimental results.
翻译:我们研究了在最大值和索引反馈下,针对最大值奖励函数的组合多臂赌博机问题。这是一种介于常见半赌博机与全赌博机反馈结构之间的新型反馈机制。我们提出了一种算法,并针对随机臂结果服从任意有限支撑分布的问题实例给出了遗憾界。该遗憾分析基于考虑一个扩展臂集合(包含臂结果的值与概率信息)并应用光滑性条件。我们的算法实现了与分布相关的$O((k/\Delta)\log(T))$遗憾界以及与分布无关的$\tilde{O}(\sqrt{T})$遗憾界,其中$k$为每轮选择的臂数,$\Delta$为依赖分布奖励差距,$T$为决策周期。令人意外的是,该遗憾界与已知的在信息更充分的半赌博机反馈下获得的界相当。我们通过实验结果验证了算法的有效性。