We examine the relationship between the mutual information between the output model and the empirical sample and the generalization of the algorithm in the context of stochastic convex optimization. Despite increasing interest in information-theoretic generalization bounds, it is uncertain if these bounds can provide insight into the exceptional performance of various learning algorithms. Our study of stochastic convex optimization reveals that, for true risk minimization, dimension-dependent mutual information is necessary. This indicates that existing information-theoretic generalization bounds fall short in capturing the generalization capabilities of algorithms like SGD and regularized ERM, which have dimension-independent sample complexity.
翻译:我们研究了在随机凸优化背景下,输出模型与经验样本之间的互信息与算法泛化性能之间的关系。尽管对信息论泛化界的研究兴趣日益增长,但这些界限能否揭示各类学习算法卓越性能的内在机理仍不确定。我们对随机凸优化的研究表明,为实现真实风险最小化,必须引入依赖于维度的互信息。这表明现有的信息论泛化界无法捕捉诸如SGD和正则化ERM等具有维度无关样本复杂度的算法的泛化能力。