The multi-armed bandit(MAB) is a classical sequential decision problem. Most work requires assumptions about the reward distribution (e.g., bounded), while practitioners may have difficulty obtaining information about these distributions to design models for their problems, especially in non-stationary MAB problems. This paper aims to design a multi-armed bandit algorithm that can be implemented without using information about the reward distribution while still achieving substantial regret upper bounds. To this end, we propose a novel algorithm alternating between greedy rule and forced exploration. Our method can be applied to Gaussian, Bernoulli and other subgaussian distributions, and its implementation does not require additional information. We employ a unified analysis method for different forced exploration strategies and provide problem-dependent regret upper bounds for stationary and piecewise-stationary settings. Furthermore, we compare our algorithm with popular bandit algorithms on different reward distributions.
翻译:多臂赌博机(MAB)是一个经典的序贯决策问题。大多数研究需要对奖励分布进行假设(例如有界性),而实践者在针对其问题设计模型时,可能难以获取这些分布的相关信息,尤其是在非平稳MAB问题中。本文旨在设计一种无需利用奖励分布信息即可实现的多臂赌博机算法,同时仍能获得显著的概率后悔上界。为此,我们提出了一种在贪婪规则与强制探索之间交替进行的新型算法。该方法可应用于高斯分布、伯努利分布及其他次高斯分布,且其实现无需额外信息。我们针对不同强制探索策略采用统一的分析方法,并针对平稳和分段平稳设定给出了问题依赖的概率后悔上界。此外,我们还将所提算法与主流赌博机算法在不同奖励分布下进行了比较。