A wide range of empirical and theoretical works have shown that overparameterisation can amplify the performance of neural networks. According to the lottery ticket hypothesis, overparameterised networks have an increased chance of containing a sub-network that is well-initialised to solve the task at hand. A more parsimonious approach, inspired by animal learning, consists in guiding the learner towards solving the task by curating the order of the examples, i.e. providing a curriculum. However, this learning strategy seems to be hardly beneficial in deep learning applications. In this work, we undertake an analytical study that connects curriculum learning and overparameterisation. In particular, we investigate their interplay in the online learning setting for a 2-layer network in the XOR-like Gaussian Mixture problem. Our results show that a high degree of overparameterisation -- while simplifying the problem -- can limit the benefit from curricula, providing a theoretical account of the ineffectiveness of curricula in deep learning.
翻译:大量实证与理论研究已表明,过参数化能够提升神经网络的性能。根据彩票假说,过参数化网络包含一个能够良好初始化以解决当前任务的子网络的概率更高。受动物学习启发,一种更简约的方法是通过精心设计样本呈现顺序(即提供课程)来引导学习者解决问题。然而,这种学习策略在深度学习应用中似乎难以显现优势。本研究通过理论分析,建立了课程学习与过参数化之间的联系。具体而言,我们在在线学习场景下,针对XOR类高斯混合问题中的双层网络,探究了两者的相互作用。研究结果表明,高度的过参数化虽然简化了问题,却可能限制课程学习带来的收益,从而为课程学习在深度学习中的低效性提供了理论解释。