Providing generalization guarantees for modern neural networks has been a crucial task in statistical learning. Recently, several studies have attempted to analyze the generalization error in such settings by using tools from fractal geometry. While these works have successfully introduced new mathematical tools to apprehend generalization, they heavily rely on a Lipschitz continuity assumption, which in general does not hold for neural networks and might make the bounds vacuous. In this work, we address this issue and prove fractal geometry-based generalization bounds without requiring any Lipschitz assumption. To achieve this goal, we build up on a classical covering argument in learning theory and introduce a data-dependent fractal dimension. Despite introducing a significant amount of technical complications, this new notion lets us control the generalization error (over either fixed or random hypothesis spaces) along with certain mutual information (MI) terms. To provide a clearer interpretation to the newly introduced MI terms, as a next step, we introduce a notion of "geometric stability" and link our bounds to the prior art. Finally, we make a rigorous connection between the proposed data-dependent dimension and topological data analysis tools, which then enables us to compute the dimension in a numerically efficient way. We support our theory with experiments conducted on various settings.
翻译:为现代神经网络提供泛化保证一直是统计学习中的关键任务。近年来,多项研究尝试利用分形几何工具分析此类设置下的泛化误差。尽管这些工作成功引入了新的数学工具来理解泛化,但它们严重依赖于Lipschitz连续性假设,而该假设通常对神经网络不成立,并可能导致泛化界失去意义。本文解决了这一问题,在无需任何Lipschitz假设的情况下,证明了基于分形几何的泛化界。为此,我们基于学习理论中的经典覆盖论证,引入了一种数据依赖的分形维数。尽管这一新概念带来了大量技术复杂性,但它使我们能够控制(固定或随机假设空间上的)泛化误差以及某些互信息(MI)项。为了更清晰地解释新引入的MI项,我们进一步提出了"几何稳定性"概念,并将我们的泛化界与先前工作建立联系。最后,我们严格论证了所提出的数据依赖维数与拓扑数据分析工具之间的关联,从而能够以数值高效的方式计算该维数。我们在多种设置下的实验验证了理论的有效性。