Linear mode-connectivity (LMC) (or lack thereof) is one of the intriguing characteristics of neural network loss landscapes. While empirically well established, it unfortunately still lacks a proper theoretical understanding. Even worse, although empirical data points are abound, a systematic study of when networks exhibit LMC is largely missing in the literature. In this work we aim to close this gap. We explore how LMC is affected by three factors: (1) architecture (sparsity, weight-sharing), (2) training strategy (optimization setup) as well as (3) the underlying dataset. We place particular emphasis on minimal but non-trivial settings, removing as much unnecessary complexity as possible. We believe that our insights can guide future theoretical works on uncovering the inner workings of LMC.
翻译:线性模式连通性(LMC)及其缺失性是神经网络损失景观的一个引人注目的特征。尽管这一现象在实证层面已得到充分验证,但遗憾的是,其仍然缺乏严格的理论理解。更糟糕的是,虽然经验数据点随处可见,但关于网络在何种条件下表现出LMC的系统性研究在文献中基本缺失。本研究旨在填补这一空白。我们探究了LMC受三个因素影响的方式:(1)架构(稀疏性、权重共享)、(2)训练策略(优化设置)以及(3)底层数据集。我们特别强调在最小但非平凡设置下的研究,尽可能去除不必要的复杂性。我们相信,这些见解能够为未来探索LMC内在机制的理论工作提供指导。