Pre-trained Large Language Models (LLMs) have shown success in a diverse set of language inference and understanding tasks. The pre-training stage of LLMs looks at a large corpus of raw textual data. The BabyLM shared task compares LLM pre-training to human language acquisition, where the number of tokens seen by 13-year-old kids is magnitudes smaller than the number of tokens seen by LLMs. In this work, we pre-train and evaluate LLMs on their ability to learn contextual word representations using roughly the same number of tokens as seen by children. We provide a strong set of baselines; with different architectures, evaluation of changes in performance across epochs, and reported pre-training metrics for the strict small and strict tracks of the task. We also try to loosely replicate the RoBERTa baseline given by the task organizers to observe the training robustness to hyperparameter selection and replicability. We provide the submission details to the strict and strict-small tracks in this report.
翻译:摘要:预训练大型语言模型在多样化的语言推理和理解任务中展现出成功。其预训练阶段需处理大量原始文本数据。BabyLM共享任务将语言模型预训练与人类语言习得过程进行对比,其中13岁儿童接触的token数量远小于语言模型所见token量。本研究使用与儿童所见大致相同数量的token,预训练并评估语言模型学习上下文词表示的能力。我们提供了一组强基线:涵盖不同架构、评估各训练周期的性能变化,并报告了任务严格小规模与严格轨道下的预训练指标。同时尝试粗略复现任务组织者提供的RoBERTa基线,以观察超参数选择对训练鲁棒性的影响及可复现性。本报告亦提交了严格轨道与严格小规模轨道的实验细节。