This paper considers the learning of logical (Boolean) functions with focus on the generalization on the unseen (GOTU) setting, a strong case of out-of-distribution generalization. This is motivated by the fact that the rich combinatorial nature of data in certain reasoning tasks (e.g., arithmetic/logic) makes representative data sampling challenging, and learning successfully under GOTU gives a first vignette of an 'extrapolating' or 'reasoning' learner. We then study how different network architectures trained by (S)GD perform under GOTU and provide both theoretical and experimental evidence that for a class of network models including instances of Transformers, random features models, and diagonal linear networks, a min-degree-interpolator (MDI) is learned on the unseen. We also provide evidence that other instances with larger learning rates or mean-field networks reach leaky MDIs. These findings lead to two implications: (1) we provide an explanation to the length generalization problem (e.g., Anil et al. 2022); (2) we introduce a curriculum learning algorithm called Degree-Curriculum that learns monomials more efficiently by incrementing supports.
翻译:本文研究逻辑(布尔)函数的学习问题,重点关注未看见数据上的泛化(Generalization on the Unseen, GOTU)这一分布外泛化强情形。该研究的动机在于,某些推理任务(如算术/逻辑)中数据丰富的组合特性使得代表性数据采样具有挑战性,而在GOTU设定下成功学习为具备“外推”或“推理”能力的学习者提供了初步雏形。我们进而研究经(S)GD训练的不同网络架构在GOTU设定下的表现,并从理论和实验两方面证明:对于包括Transformer实例、随机特征模型和对角线性网络在内的某类网络模型,未看见数据上学习到的是最小度数插值器(Min-Degree-Interpolator, MDI)。我们还提供证据表明,采用较大学习率的其他实例或平均场网络会趋近于泄露型MDI。这些发现带来两点启示:(1)为长度泛化问题(如Anil等人2022年研究)提供了一种解释;(2)提出一种名为“度数课程”(Degree-Curriculum)的课程学习算法,该算法通过逐步增加支撑集更高效地学习单项式。