This paper details the work of the University of Groningen for the BabyLM Challenge. We follow the idea that, like babies, language models should be introduced to simpler concepts first and build off of that knowledge to understand more complex concepts. We examine this strategy of simple-then-complex through a variety of lenses, namely context size, vocabulary, and overall linguistic complexity of the data. We find that only one, context size, is truly beneficial to training a language model. However this simple change to context size gives us improvements of 2 points on average on (Super)GLUE tasks, 1 point on MSGS tasks, and 12\% on average on BLiMP tasks. Our context-limited model outperforms the baseline that was trained on 10$\times$ the amount of data.
翻译:本文详细介绍了格罗宁根大学在BabyLM挑战赛中的工作。我们遵循以下理念:与婴儿类似,语言模型应首先接触简单概念,并以此为基础逐步理解更复杂的概念。我们从多个维度考察了这种从简单到复杂的策略,包括上下文规模、词汇量以及数据的整体语言复杂度。研究发现,只有上下文规模这一维度对语言模型训练真正有益。然而,仅通过调整上下文规模这一简单变化,模型在(Super)GLUE任务上平均提升2分,在MSGS任务上提升1分,在BLiMP任务上平均提升12%。我们的上下文受限模型性能超越了基于10倍数据量训练的基线模型。