In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexities necessarily arise for feature learning in the presence of computational-statistical gaps. We begin by considering offline sparse parity learning, a supervised classification problem which admits a statistical query lower bound for gradient-based training of a multilayer perceptron. This lower bound can be interpreted as a multi-resource tradeoff frontier: successful learning can only occur if one is sufficiently rich (large model), knowledgeable (large dataset), patient (many training iterations), or lucky (many random guesses). We show, theoretically and experimentally, that sparse initialization and increasing network width yield significant improvements in sample efficiency in this setting. Here, width plays the role of parallel search: it amplifies the probability of finding "lottery ticket" neurons, which learn sparse features more sample-efficiently. Finally, we show that the synthetic sparse parity task can be useful as a proxy for real problems requiring axis-aligned feature learning. We demonstrate improved sample efficiency on tabular classification benchmarks by using wide, sparsely-initialized MLP models; these networks sometimes outperform tuned random forests.
翻译:在现代深度学习中,算法选择(如宽度、深度和学习率)已知会调节微妙的资源权衡。本文研究了在计算-统计差距存在的情况下,这些复杂性如何必然出现在特征学习中。我们首先考虑离线稀疏奇偶学习问题,这是一个有监督分类问题,其对多层感知器的梯度训练存在统计查询下界。该下界可解释为多资源权衡前沿:成功学习只有在资源足够丰富(大模型)、知识足够丰富(大数据集)、训练足够耐心(多次训练迭代)或运气足够好(多次随机猜测)时才能实现。我们从理论和实验上证明,稀疏初始化和增加网络宽度能在该设置下显著提升样本效率。此处宽度扮演了并行搜索的角色:它放大了找到"彩票神经元"的概率,这些神经元能以更高样本效率学习稀疏特征。最后,我们展示合成稀疏奇偶任务可作为需要轴对齐特征学习的实际问题的有效代理。通过使用宽度大、稀疏初始化的多层感知器模型,我们在表格分类基准上展示了改进的样本效率;这些网络有时能超越调优后的随机森林。