To model behavioral and neural correlates of language comprehension in naturalistic environments researchers have turned to broad-coverage tools from natural-language processing and machine learning. Where syntactic structure is explicitly modeled, prior work has relied predominantly on context-free grammars (CFG), yet such formalisms are not sufficiently expressive for human languages. Combinatory Categorial Grammars (CCGs) are sufficiently expressive directly compositional models of grammar with flexible constituency that affords incremental interpretation. In this work we evaluate whether a more expressive CCG provides a better model than a CFG for human neural signals collected with fMRI while participants listen to an audiobook story. We further test between variants of CCG that differ in how they handle optional adjuncts. These evaluations are carried out against a baseline that includes estimates of next-word predictability from a Transformer neural network language model. Such a comparison reveals unique contributions of CCG structure-building predominantly in the left posterior temporal lobe: CCG-derived measures offer a superior fit to neural signals compared to those derived from a CFG. These effects are spatially distinct from bilateral superior temporal effects that are unique to predictability. Neural effects for structure-building are thus separable from predictability during naturalistic listening, and those effects are best characterized by a grammar whose expressive power is motivated on independent linguistic grounds.
翻译:为了对自然环境中语言理解的行为和神经关联进行建模,研究人员已转向自然语言处理和机器学习中的广覆盖工具。在明确建模句法结构时,先前的工作主要依赖于上下文无关文法(CFG),然而这种形式体系不足以表达人类语言的复杂性。组合范畴文法(CCG)是一种具有充分表达能力且直接组合的语法模型,其灵活的短语结构支持增量式解读。本研究评估了更具表达力的CCG是否比CFG更能匹配参与者聆听有声书故事时通过fMRI采集的人脑神经信号。我们进一步测试了CCG中处理可选附加语的不同变体。这些评估以Transformer神经网络语言模型对下一词可预测性的估计为基准进行。此类比较揭示了CCG结构构建在左后颞叶的独特贡献:与CFG相比,CCG衍生指标对神经信号的拟合效果更优。这些效应在空间分布上区别于双侧颞上回中与可预测性相关的独特效应。因此,自然聆听过程中结构构建的神经效应可与可预测性分离,且这些效应最能由具有独立语言学动机表达能力的语法所刻画。