Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal representations of Large Language Models (LLMs) support their ability to parse text when predicting the next word, while representing semantic notions independently of surface form. Yet, which data statistics make these feats possible, and how much data is required, remain largely unknown. Probabilistic context-free grammars (PCFGs) provide a tractable testbed for studying these questions. However, prior work has focused either on the post-hoc characterization of the parsing-like algorithms used by trained networks; or on the learnability of PCFGs with fixed syntax, where parsing is unnecessary. Here, we (i) introduce a tunable class of PCFGs in which both the degree of ambiguity and the correlation structure across scales can be controlled; (ii) provide a learning mechanism -- an inference algorithm inspired by the structure of deep convolutional networks -- that links learnability and sample complexity to specific language statistics; and (iii) validate our predictions empirically across deep convolutional and transformer-based architectures. Overall, we propose a unifying framework where correlations at different scales lift local ambiguities, enabling the emergence of hierarchical representations of the data.
翻译:理解如何仅从句子中学习语言结构是认知科学和机器学习领域的核心问题。对大型语言模型内部表征的研究表明,它们在预测下一个词时能够解析文本,同时独立于表面形式表达语义概念。然而,哪些数据统计特征使这些能力成为可能,以及需要多少数据,在很大程度上仍不明确。概率上下文无关文法为研究这些问题提供了一个可操作测试平台。然而,先前的工作要么侧重于对训练网络所使用的类解析算法的后验表征,要么侧重于固定语法(无需解析)下PCFG的可学习性。在此,我们(i)引入一类可调参的PCFG,其中歧义程度和跨尺度的相关性结构均可被控制;(ii)提供一种学习机制——受深层卷积网络结构启发的推理算法——将可学习性和样本复杂度与特定语言统计特征相关联;以及(iii)在深层卷积和基于Transformer的架构上实证验证我们的预测。总体而言,我们提出了一个统一框架,其中不同尺度的相关性消除了局部歧义,从而促进了数据层次化表征的出现。