Multi-armed bandits are one of the theoretical pillars of reinforcement learning. Recently, the investigation of quantum algorithms for multi-armed bandit problems was started, and it was found that a quadratic speed-up (in query complexity) is possible when the arms and the randomness of the rewards of the arms can be queried in superposition. Here we introduce further bandit models where we only have limited access to the randomness of the rewards, but we can still query the arms in superposition. We show that then the query complexity is the same as for classical algorithms. This generalizes the prior result that no speed-up is possible for unstructured search when the oracle has positive failure probability.
翻译:多臂老虎机是强化学习的理论支柱之一。近年来,针对多臂老虎机问题的量子算法研究已经展开,研究发现当臂以及臂的奖励随机性可以叠加态查询时,可实现(查询复杂度上的)二次加速。在此,我们引入进一步的虎机模型,其中对奖励随机性的访问有限,但仍可对臂进行叠加态查询。我们证明,此时查询复杂度与经典算法相同。这一结果推广了先前关于非结构化搜索中当预言存在正失败概率时无法实现加速的结论。