We prove that a classic sub-Gaussian mixture proposed by Robbins in a stochastic setting actually satisfies a path-wise (deterministic) regret bound. For every path in a natural ``Ville event'' $\mathcal E_α$, this regret till time $T$ is bounded by $\ln^2(1/α)/V_T + \ln (1/α) + \ln \ln V_T$ up to universal constants, where $V_T$ is a nonnegative, nondecreasing, cumulative variance process. (The bound reduces to $\ln(1/α) + \ln \ln V_T$ if $V_T \geq \ln(1/α)$.) If the data were stochastic, then one can show that $\mathcal E_α$ has probability at least $1-α$ under a wide class of distributions (eg: sub-Gaussian, symmetric, variance-bounded, etc.). In fact, we show that on the Ville event $\mathcal E_0$ of probability one, the regret on every path in $\mathcal E_0$ is eventually bounded by $\ln \ln V_T$ (up to constants). We explain how this work helps bridge the world of adversarial online learning (which usually deals with regret bounds for bounded data), with game-theoretic statistics (which can handle unbounded data, albeit using stochastic assumptions). In short, conditional regret bounds serve as a bridge between stochastic and adversarial betting.
翻译:我们证明,Robbins在随机背景下提出的经典亚高斯混合实际上满足路径(确定性)后悔界。对于自然“Ville事件” $\mathcal E_\alpha$ 中的每条路径,到时间 $T$ 为止的后悔在通用常数意义下由 $\ln^2(1/\alpha)/V_T + \ln(1/\alpha) + \ln\ln V_T$ 界定(若 $V_T \geq \ln(1/\alpha)$,则该界简化为 $\ln(1/\alpha) + \ln\ln V_T$),其中 $V_T$ 为非负、非降的累积方差过程。若数据具有随机性,则可证明在广泛分布类(例如:亚高斯、对称、有界方差等)下,$\mathcal E_\alpha$ 的概率至少为 $1-\alpha$。事实上,我们表明在概率为1的Ville事件 $\mathcal E_0$ 上,$\mathcal E_0$ 中每条路径的后悔最终由 $\ln\ln V_T$(至多常数项)界定。本文阐释了该工作如何架起对抗在线学习(通常处理有界数据的后悔界)与博弈统计(虽然使用随机假设,但能处理无界数据)之间的桥梁。简言之,条件后悔界可作为随机博弈与对抗博弈之间的纽带。