Word frequency is a strong predictor in most lexical processing tasks. Thus, any model of word recognition needs to account for how word frequency effects arise. The Discriminative Lexicon Model (DLM; Baayen et al., 2018a, 2019) models lexical processing with linear mappings between words' forms and their meanings. So far, the mappings can either be obtained incrementally via error-driven learning, a computationally expensive process able to capture frequency effects, or in an efficient, but frequency-agnostic solution modelling the theoretical endstate of learning (EL) where all words are learned optimally. In this study we show how an efficient, yet frequency-informed mapping between form and meaning can be obtained (Frequency-informed learning; FIL). We find that FIL well approximates an incremental solution while being computationally much cheaper. FIL shows a relatively low type- and high token-accuracy, demonstrating that the model is able to process most word tokens encountered by speakers in daily life correctly. We use FIL to model reaction times in the Dutch Lexicon Project (Keuleers et al., 2010) and find that FIL predicts well the S-shaped relationship between frequency and the mean of reaction times but underestimates the variance of reaction times for low frequency words. FIL is also better able to account for priming effects in an auditory lexical decision task in Mandarin Chinese (Lee, 2007), compared to EL. Finally, we used ordered data from CHILDES (Brown, 1973; Demuth et al., 2006) to compare mappings obtained with FIL and incremental learning. The mappings are highly correlated, but with FIL some nuances based on word ordering effects are lost. Our results show how frequency effects in a learning model can be simulated efficiently, and raise questions about how to best account for low-frequency words in cognitive models.
翻译:词频是大多数词汇处理任务中的强预测因子。因此,任何词汇识别模型都需要解释词频效应是如何产生的。判别性词汇模型(DLM;Baayen等人,2018a,2019)通过单词形式与其意义之间的线性映射来模拟词汇处理。目前,这些映射可以通过基于误差驱动的增量学习(一种计算成本高昂但能捕捉频率效应的方法)获得,也可以通过一种高效但忽略频率的解来建模学习的理论终态(EL),其中所有单词都得到最优学习。在本研究中,我们展示了如何获得一种既高效又考虑频率的形式与意义之间的映射(频率知情学习;FIL)。我们发现,FIL能很好地逼近增量解,同时计算成本显著降低。FIL表现出相对较低的类型准确率和较高的词例准确率,表明该模型能正确处理说话者在日常生活中遇到的大多数词例。我们使用FIL对荷兰语词汇项目(Keuleers等人,2010)中的反应时间进行建模,发现FIL能较好地预测频率与反应时间均值之间的S形关系,但低估了低频单词反应时间的方差。与EL相比,FIL在普通话听觉词汇决策任务(Lee,2007)中也能更好地解释启动效应。最后,我们使用来自CHILDES的有序数据(Brown,1973;Demuth等人,2006)来比较FIL与增量学习获得的映射。两种映射高度相关,但FIL丢失了一些基于词序效应的细微差别。我们的结果展示了如何在学习模型中高效模拟频率效应,并提出了在认知模型中如何最好地解释低频单词的问题。