Traditionally, large language models have been either trained on general web crawls or domain-specific data. However, recent successes of generative large language models, have shed light on the benefits of cross-domain datasets. To examine the significance of prioritizing data diversity over quality, we present a German dataset comprising texts from five domains, along with another dataset aimed at containing high-quality data. Through training a series of models ranging between 122M and 750M parameters on both datasets, we conduct a comprehensive benchmark on multiple downstream tasks. Our findings demonstrate that the models trained on the cross-domain dataset outperform those trained on quality data alone, leading to improvements up to $4.45\%$ over the previous state-of-the-art. The models are available at https://huggingface.co/ikim-uk-essen
翻译:传统上,大型语言模型要么基于通用网络爬取数据,要么基于特定领域数据进行训练。然而,近期生成式大型语言模型取得的成功揭示了跨领域数据集的优势。为探究优先考虑数据多样性而非数据质量的重要性,我们构建了一个包含五个领域文本的德语数据集,以及另一个旨在包含高质量数据的数据集。通过在两个数据集上训练一系列参数量介于1.22亿至7.5亿之间的模型,我们对多个下游任务进行了全面基准测试。实验结果表明,基于跨领域数据集训练的模型优于仅使用高质量数据训练的模型,相较于先前最优水平提升最高达$4.45\%$。相关模型已发布于https://huggingface.co/ikim-uk-essen