We find limits to the Transformer architecture for language modeling and show it has a universal prediction property in an information-theoretic sense. We further analyze performance in non-asymptotic data regimes to understand the role of various components of the Transformer architecture, especially in the context of data-efficient training. We validate our theoretical analysis with experiments on both synthetic and real datasets.
翻译:我们发现了Transformer架构在语言建模中的局限性,并从信息论角度证明其具有通用预测性质。进一步地,我们分析了非渐近数据机制下的性能表现,以理解Transformer架构各组件的作用,特别是在数据高效训练场景下的影响。我们通过合成数据集和真实数据集的实验验证了理论分析的正确性。