The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conjecturing a positive 'quality-power' relationship, in which language models' (LMs') fit to psychometric data continues to improve as their ability to predict words in context increases. This is important because it might suggest that elements of LLM architecture reflect the architecture of the human sentence processing faculty, and that any inadequacies in predicting human reading time and brain imaging data may be attributed to insufficient model complexity, which recedes as larger models become available. But recent studies have shown this scaling inverts after a point, as LMs become excessively large and accurate, when information-theoretic surprisal is used as a predictor. Other studies propose the use of entire vectors from differently sized LLMs, still showing positive scaling, casting doubt on the value of surprisal as a predictor, but do not control for dimensionality expansion using untrained LLMs with more than 1.6B parameters. This study evaluates scaling of LLM vector predictors controlled using untrained LLMs with up to 66B parameters. Results show that inverse scaling obtains, and moreover the contribution of trained LMs over corresponding untrained LMs drops to zero at around a few billion parameters on most datasets.
翻译:暂无翻译