Recent research has shown that filtering massive English web corpora into high-quality subsets significantly improves training efficiency. However, for high-resource non-English languages like German, French, or Japanese, aggressive filtering creates a strategic dilemma: should practitioners prioritize diversity by training once on large amounts of lightly filtered web data, or prioritize quality by strictly filtering for a high-quality core and repeating it over multiple epochs? We investigate this trade-off for German by constructing hierarchical quality filters applied to 500M web documents, comparing multi-epoch training on the filtered subsets against single-pass training on a diverse corpus. Our experiments across multiple model scales and token budgets show that repeating high-quality data consistently outperforms single-pass training on larger, less filtered sets. Notably, the performance gap persists even after 7 epochs. Our findings suggest that for non-English LLMs, semantic concentration through quality filtering offers a more viable path to efficient language modeling than simply maximizing unique data volume. We release our German language models (called Boldt), as well as our cleaned evaluation benchmarks to the research community. Our experiments indicate that they achieve state-of-the-art results despite training on 10-360x fewer tokens than comparable models.
翻译:近期研究表明,从大规模英文网页语料中过滤出高质量子集能显著提升训练效率。然而,针对德语、法语、日语等资源丰富的非英语语言,激进的数据过滤策略会引发战略性两难:从业者应优先追求多样性(对大量轻度过滤的网络数据进行单次训练),还是优先追求质量(严格筛选高质核心数据并在多个训练周期中重复使用)?我们通过构建分层质量过滤器对5亿篇网页文档进行处理,系统探究了德语场景下的这一权衡,并比较了在过滤子集上进行多周期训练与在多样化语料上进行单次训练的效果。跨多个模型规模和令牌预算的实验表明,重复使用高质量数据始终优于在更大规模但过滤不充分的集合上进行单次训练。值得注意的是,即便经过7个训练周期,这一性能差距依然存在。研究结果表明,对于非英语大语言模型而言,通过质量过滤实现语义聚焦比单纯最大化独特数据量更能实现高效语言建模。我们向研究社区发布了德语语言模型(命名为Boldt)及清洗后的评估基准。实验表明,尽管训练所用的令牌数仅为同类模型的1/10至1/360,该模型仍取得了当前最优性能。