Thompson sampling (TS) is widely used for stochastic multi-armed bandits, yet its inferential properties under adaptive data collection are subtle. Classical asymptotic theory for sample means can fail because arm-specific sample sizes are random and coupled with the rewards through the action-selection rule. We study adaptive inference for Thompson sampling with Gaussian randomized indices in $K$-armed stochastic bandits with independent sub-Gaussian reward noises, and identify \emph{optimism} as a key mechanism for restoring \emph{stability}, meaning that each arm's pull count concentrates around a deterministic scale. This stability yields asymptotically valid Wald inference despite adaptive sampling. First, we prove that variance-inflated TS is stable for any $K \ge 2$, including the challenging regime where multiple arms are optimal, with asymptotically uniform allocation over optimal arms and sharp logarithmic pull-count asymptotics for suboptimal arms. This resolves the $K$-armed extension question raised by \citet{halder2025stable}, using new winner-map and Lyapunov-drift techniques to control allocation among multiple optimal arms. Second, we analyze an alternative optimistic modification that keeps the Gaussian index variance unchanged but adds an explicit mean bonus to the index center, and establish a similar stability conclusion. In summary, suitably implemented optimism stabilizes Thompson sampling and enables asymptotically valid Wald inference in multi-armed bandits, while incurring only a mild additional regret cost.
翻译:汤普森采样(TS)广泛用于随机多臂老虎机问题,但其在自适应数据收集下的推断性质较为微妙。样本均值的经典渐近理论可能失效,因为臂特定的样本量是随机的,并通过动作选择规则与奖励关联。我们研究了带有高斯随机索引的汤普森采样在$K$臂随机老虎机(具有独立次高斯奖励噪声)中的自适应推断,并识别出\emph{乐观主义}是恢复\emph{稳定性}的关键机制,即每臂的拉动次数集中在确定性尺度附近。尽管存在自适应采样,这种稳定性仍能产生渐近有效的Wald推断。首先,我们证明对于任意$K \ge 2$,方差膨胀的TS是稳定的,包括多个臂最优的挑战性情况,此时最优臂的分配渐近均匀,而次优臂的拉动次数具有尖锐的对数渐近性。这解决了\citet{halder2025stable}提出的$K$臂扩展问题,我们利用新的获胜者映射和Lyapunov漂移技术来控制多个最优臂之间的分配。其次,我们分析了一种替代的乐观修改,该方法保持高斯索引方差不变,但向索引中心添加显式的均值奖励,并建立类似的稳定性结论。总之,适当实施的乐观主义稳定了汤普森采样,并在多臂老虎机中实现了渐近有效的Wald推断,同时仅带来轻微的额外遗憾成本。