We present our proposed solution to the BabyLM challenge [arXiv:2301.11796], whose goal was to improve the sample efficiency of language models. We trained an ensemble consisting of a GPT-2 and small LLaMA models on the developmentally-plausible, 10M-word BabyLM dataset, then distilled it into a small, 58M-parameter LLaMA model, which exceeds in performance both of its teachers as well as a similar model trained without distillation. This suggests that distillation can not only retain the full performance of the teacher model when the latter is trained on a sufficiently small dataset; it can exceed it, and lead to significantly better performance than direct training.
翻译:我们提出了针对BabyLM挑战[arXiv:2301.11796]的解决方案,该挑战旨在提升语言模型的样本效率。我们在符合发展规律、包含1000万词的BabyLM数据集上训练了一个由GPT-2和小型LLaMA模型组成的集成模型,随后将其蒸馏为参数量仅5800万的小型LLaMA模型。该模型在性能上不仅超越了其所有教师模型,还优于未经蒸馏训练的同类模型。这表明:当教师模型在足够小的数据集上训练时,蒸馏不仅能完整保留其性能,甚至能超越教师模型,并带来显著优于直接训练的效果。