Current trends to pre-train capable Large Language Models (LLMs) mostly focus on scaling of model and dataset size. However, the quality of pre-training data is an important factor for training powerful LLMs, yet it is a nebulous concept that has not been fully characterized. Therefore, we use the recently proposed Task2Vec diversity coefficient to ground and understand formal aspects of data quality, to go beyond scale alone. Specifically, we measure the diversity coefficient of publicly available pre-training datasets to demonstrate that their formal diversity is high when compared to theoretical lower and upper bounds. In addition, to build confidence in the diversity coefficient, we conduct interpretability experiments and find that the coefficient aligns with intuitive properties of diversity, e.g., it increases as the number of latent concepts increases. We conclude the diversity coefficient is reliable, show it's high for publicly available LLM datasets, and conjecture it can be used to build useful diverse datasets for LLMs.
翻译:当前训练强大大型语言模型(LLMs)的主流趋势主要聚焦于模型与数据集规模的扩展。然而,预训练数据质量对于训练高性能LLMs至关重要,但这一概念仍模糊不清且尚未被充分定义。为此,我们采用近期提出的Task2Vec多样性系数作为量化框架,通过超越单纯规模视角来理解数据的形式化特征。具体而言,我们测量了公开可用预训练数据集的多样性系数,发现相较于理论下界与上界,这些数据的多样性处于较高水平。此外,为验证该系数的可靠性,我们开展了可解释性实验,结果表明多样性系数与人类直觉中的多样性属性高度一致——例如,当隐式概念数量增加时系数值随之上升。我们得出结论:多样性系数具有可靠的可解释性,证实公开LLM数据集具有高多样性,并推测该指标可用于构建适用于LLMs的高质量多样化数据集。