In recent years, significant attention in deep learning theory has been devoted to analyzing when models that interpolate their training data can still generalize well to unseen examples. Many insights have been gained from studying models with multiple layers of Gaussian random features, for which one can compute precise generalization asymptotics. However, few works have considered the effect of weight anisotropy; most assume that the random features are generated using independent and identically distributed Gaussian weights, and allow only for structure in the input data. Here, we use the replica trick from statistical physics to derive learning curves for models with many layers of structured Gaussian features. We show that allowing correlations between the rows of the first layer of features can aid generalization, while structure in later layers is generally detrimental. Our results shed light on how weight structure affects generalization in a simple class of solvable models.
翻译:近年来,深度学习理论领域的大量研究集中于分析:那些能完美拟合训练数据的模型,为何仍能对未见样本实现良好的泛化。通过研究多层高斯随机特征模型(这类模型可精确计算泛化渐近性能),研究者已获得诸多洞见。然而,现有工作较少考虑权重各向异性的影响——多数研究假设随机特征由独立同分布的高斯权重生成,仅允许输入数据存在结构。本文采用统计物理中的副本技巧,推导了含有多层结构化高斯特征模型的学习曲线。研究表明,允许第一层特征行之间存在相关性有助于泛化,而后层结构通常具有负面效应。我们的结果揭示了在简单可解模型类别中,权重结构如何影响泛化性能。