Representations from transformer-based unidirectional language models are known to be effective at predicting brain responses to natural language. However, most studies comparing language models to brains have used GPT-2 or similarly sized language models. Here we tested whether larger open-source models such as those from the OPT and LLaMA families are better at predicting brain responses recorded using fMRI. Mirroring scaling results from other contexts, we found that brain prediction performance scales log-linearly with model size from 125M to 30B parameter models, with ~15% increased encoding performance as measured by correlation with a held-out test set across 3 subjects. Similar log-linear behavior was observed when scaling the size of the fMRI training set. We also characterized scaling for acoustic encoding models that use HuBERT, WavLM, and Whisper, and we found comparable improvements with model size. A noise ceiling analysis of these large, high-performance encoding models showed that performance is nearing the theoretical maximum for brain areas such as the precuneus and higher auditory cortex. These results suggest that increasing scale in both models and data will yield incredibly effective models of language processing in the brain, enabling better scientific understanding as well as applications such as decoding.
翻译:基于Transformer的单向语言模型的表征已被证明能有效预测大脑对自然语言的响应。然而,多数将语言模型与大脑进行比较的研究仍采用GPT-2或同等规模的语言模型。本研究测试了更大规模的开源模型(如OPT和LLaMA系列)是否能更好预测基于fMRI记录的大脑响应。与其他领域的缩放结果一致,我们发现大脑预测性能随模型规模(1.25亿至300亿参数)呈对数线性增长,与三个受试者保留测试集的相关系数相比,编码性能提升约15%。当缩放fMRI训练集的规模时,也观察到类似的对数线性行为。我们还刻画了基于HuBERT、WavLM和Whisper的声学编码模型的缩放规律,发现模型规模同样带来可比性的性能提升。对这些大规模高性能编码模型的噪声上限分析表明,在楔前叶和高级听觉皮层等脑区,其性能已接近理论最大值。这些结果表明,通过扩大模型和数据规模,将能构建出极为有效的大脑语言处理模型,既促进科学理解,也推动解码等应用的发展。