Autoregressive processes naturally arise in a large variety of real-world scenarios, including stock markets, sales forecasting, weather prediction, advertising, and pricing. When facing a sequential decision-making problem in such a context, the temporal dependence between consecutive observations should be properly accounted for guaranteeing convergence to the optimal policy. In this work, we propose a novel online learning setting, namely, Autoregressive Bandits (ARBs), in which the observed reward is governed by an autoregressive process of order $k$, whose parameters depend on the chosen action. We show that, under mild assumptions on the reward process, the optimal policy can be conveniently computed. Then, we devise a new optimistic regret minimization algorithm, namely, AutoRegressive Upper Confidence Bound (AR-UCB), that suffers sublinear regret of order $\widetilde{\mathcal{O}} \left( \frac{(k+1)^{3/2}\sqrt{nT}}{(1-\Gamma)^2}\right)$, where $T$ is the optimization horizon, $n$ is the number of actions, and $\Gamma < 1$ is a stability index of the process. Finally, we empirically validate our algorithm, illustrating its advantages w.r.t. bandit baselines and its robustness to misspecification of key parameters.
翻译:自回归过程广泛存在于股票市场、销售预测、天气预报、广告和定价等众多现实场景中。在此类背景下进行序贯决策时,需恰当处理连续观测值之间的时间依赖性,以确保收敛到最优策略。本文提出一种新颖的在线学习框架——自回归赌博机(Autoregressive Bandits, ARBs),其中观测到的奖励受阶数为$k$的自回归过程支配,且过程参数取决于所选动作。我们证明了在奖励过程的温和假设下,可便捷计算最优策略。随后,我们设计了一种新的乐观遗憾最小化算法——自回归上置信界(AutoRegressive Upper Confidence Bound, AR-UCB),该算法可实现阶为$\widetilde{\mathcal{O}} \left( \frac{(k+1)^{3/2}\sqrt{nT}}{(1-\Gamma)^2}\right)$的次线性遗憾,其中$T$为优化时域,$n$为动作数量,$\Gamma < 1$为过程稳定性指标。最后,我们通过实验验证了该算法的有效性,展示了其相对于赌博机基线的优势以及对关键参数误设的鲁棒性。