We present ToddlerBERTa, a BabyBERTa-like language model, exploring its capabilities through five different models with varied hyperparameters. Evaluating on BLiMP, SuperGLUE, MSGS, and a Supplement benchmark from the BabyLM challenge, we find that smaller models can excel in specific tasks, while larger models perform well with substantial data. Despite training on a smaller dataset, ToddlerBERTa demonstrates commendable performance, rivalling the state-of-the-art RoBERTa-base. The model showcases robust language understanding, even with single-sentence pretraining, and competes with baselines that leverage broader contextual information. Our work provides insights into hyperparameter choices, and data utilization, contributing to the advancement of language models.
翻译:我们提出ToddlerBERTa,一种类BabyBERTa的语言模型,通过五种不同超参数的模型探索其能力。在BLiMP、SuperGLUE、MSGS以及BabyLM挑战赛的补充基准测试上评估后,我们发现较小模型在特定任务中表现优异,而较大模型在充足数据下表现良好。尽管仅在小规模数据集上训练,ToddlerBERTa仍展现出卓越性能,与当前最优的RoBERTa-base模型相抗衡。该模型即使仅使用单句预训练,也展现出强大的语言理解能力,并能与利用更广泛上下文信息的基线模型竞争。我们的工作为超参数选择与数据利用提供了洞见,推动了语言模型的发展。