What can an agent learn in a stochastic Multi-Armed Bandit (MAB) problem from a dataset that contains just a single sample for each arm? Surprisingly, in this work, we demonstrate that even in such a data-starved setting it may still be possible to find a policy competitive with the optimal one. This paves the way to reliable decision-making in settings where critical decisions must be made by relying only on a handful of samples. Our analysis reveals that \emph{stochastic policies can be substantially better} than deterministic ones for offline decision-making. Focusing on offline multi-armed bandits, we design an algorithm called Trust Region of Uncertainty for Stochastic policy enhancemenT (TRUST) which is quite different from the predominant value-based lower confidence bound approach. Its design is enabled by localization laws, critical radii, and relative pessimism. We prove that its sample complexity is comparable to that of LCB on minimax problems while being substantially lower on problems with very few samples. Finally, we consider an application to offline reinforcement learning in the special case where the logging policies are known.
翻译:在随机多臂老虎机(MAB)问题中,如果数据集仅包含每个臂的一个样本,智能体能学到什么?令人惊讶的是,本文证明即使在如此数据匮乏的环境中,仍有可能找到与最优策略相竞争的策略。这为在仅依赖少量样本做出关键决策的场景中实现可靠决策奠定了基础。我们的分析揭示了:随机策略在离线决策中可能显著优于确定性策略。针对离线多臂老虎机问题,我们设计了一种名为“随机策略增强的信任不确定性域(TRUST)”算法,该算法与主流的基于值函数的下置信界方法截然不同。其设计依赖于局部化法则、临界半径和相对悲观主义原则。我们证明其样本复杂度在极小化极大问题上与LCB相当,而在样本极少的场景中则显著更低。最后,我们将其应用于日志策略已知的离线强化学习特例场景。