In the past 20 years, artificial neural networks have become dominant in various areas, continually growing in scale. However, the current analysis of large models has mainly focused on functionality, overlooking the influence of scale differences on their properties. To address this, we propose the concept of Emergence Learning, which emphasizes the significance of scale. By studying models of different scales, we have identified a key factor in achieving higher performance in large models: the decrease of monosemantic neurons. Building on this insight, we propose a proactive approach to inhibit monosemanticity for improved performance. Our solution involves a two-phase process that includes monosemantic neuron detection and inhibition, supported by theoretical analysis. Experimental results on various tasks and neural networks demonstrate the effectiveness of our proposed method. Following the idea of Emergence Learning, though drawing inspiration from scaling phenomena, the applicability of our method is not restricted to large scale alone. Therefore, the experiment is self-contained. However, extending this research to very large-scale datasets is appealing yet impossible for research departments due to limited resources. We are delighted to share the first co-authorship and eagerly await collaboration from any AI company before submission.
翻译:在过去20年中,人工神经网络在多个领域占据主导地位,并持续扩大规模。然而,当前对大型模型的分析主要聚焦于功能性,忽视了规模差异对其特性的影响。为此,我们提出"涌现学习"概念,强调规模的重要性。通过研究不同规模的模型,我们确定了大型模型实现更高性能的关键因素:单语义神经元的减少。基于这一发现,我们提出一种主动抑制单语义性以提升性能的方法。该方案包含单语义神经元检测与抑制两个阶段,并得到理论分析支持。在多种任务和神经网络上的实验结果表明了所提方法的有效性。遵循涌现学习的思想,尽管受缩放现象的启发,但本方法的适用性并不局限于大规模场景,因此实验具有自洽性。然而,将研究扩展到超大规模数据集虽具吸引力,但受限于研究部门的资源条件尚无法实现。我们欣然分享第一作者身份,并热切期待在投稿前与任何人工智能公司的合作。