We derive generic information-theoretic and PAC-Bayesian generalization bounds involving an arbitrary convex comparator function, which measures the discrepancy between the training and population loss. The bounds hold under the assumption that the cumulant-generating function (CGF) of the comparator is upper-bounded by the corresponding CGF within a family of bounding distributions. We show that the tightest possible bound is obtained with the comparator being the convex conjugate of the CGF of the bounding distribution, also known as the Cram\'er function. This conclusion applies more broadly to generalization bounds with a similar structure. This confirms the near-optimality of known bounds for bounded and sub-Gaussian losses and leads to novel bounds under other bounding distributions.
翻译:我们推导了通用的信息论与PAC-贝叶斯泛化边界,其中涉及任意凸比较算子,该算子用于度量训练损失与总体损失之间的差异。这些边界成立的假设条件是:比较算子的累积生成函数(CGF)被某个边界分布族内对应CGF的上界所控制。研究表明,当比较算子取为边界分布CGF的凸共轭(即克拉默函数)时,可获得最紧致的边界。这一结论更广泛地适用于具有类似结构的泛化边界。这不仅验证了有界损失与次高斯损失已知边界的近最优性,还导出了其他边界分布下的新边界。