While language models achieve impressive results, their learning dynamics are far from understood. Many domains of interest -- such as natural language syntax, coding languages, arithmetic -- are captured by context-free grammars (CFGs). In this work, we extend prior work on neural language modeling of CFGs in a novel direction: how language modeling behaves with respect to CFG substructure, namely subgrammars. We define subgrammars, and prove a set of fundamental theorems connecting language modeling and subgrammars. We show that language modeling loss recurses linearly over its top-level subgrammars; applied recursively, the loss decomposes into losses for "irreducible" subgrammars. Under additional assumptions, and empirically, parametrized models learn subgrammars in parallel, unlike children who first master simple substructures. We find that subgrammar pretraining can improve final performance, but only for tiny models relative to the grammar, while alignment analyses show that pretraining consistently leads to internal representations that better reflect the grammar's substructure.
翻译:尽管语言模型取得了显著成果,但其学习动态仍远未得到充分理解。许多相关领域——如自然语言句法、编程语言、算术——均由上下文无关文法(CFG)描述。本研究在现有神经语言建模CFG工作的基础上,开辟了一个新方向:探究语言建模行为与CFG子结构(即子文法)的关系。我们定义了子文法概念,并证明了一系列连接语言建模与子文法的基本定理。研究表明,语言建模的损失函数在其顶层子文法上呈线性递归;通过递归分解,该损失可进一步拆解为"不可约"子文法的损失。在附加假设及实验验证下,参数化模型会并行学习子文法,这与儿童先掌握简单子结构的学习模式不同。我们发现,子文法预训练能提升最终性能,但仅对相对于文法规模较小的模型有效;而对齐分析则表明,预训练一贯能使模型内部表征更准确地反映文法的子结构。