Synthetic data generation through document rewriting has emerged as a promising technique for improving language model pretraining, yet most studies focus on English and do not systematically control for the quality of the source data being rewritten. We present a controlled study of how synthetic rewriting interacts with source data quality in the context of Portuguese continued pretraining. Starting from ClassiCC-PT, a Portuguese corpus annotated with STEM and Educational quality scores, we construct two 10B-token subsets at different quality levels and rewrite each into four styles using a 7B instruction-tuned model, producing approximately 40B tokens of synthetic data per condition. We train two English-centric base models (1.1B and 7B parameters) on each condition and evaluate on PoETa V2, a comprehensive 44-task Portuguese benchmark. At the 7B scale, rewriting high-quality data yields a +3.4 NPM gain over the same data unmodified, while rewriting low-quality data provides only +0.5 NPM. At the 1.1B scale, this interaction is weaker, with unmodified low-quality data performing comparably to rewritten high-quality data. Our results demonstrate that synthetic rewriting acts primarily as a quality multiplier rather than a substitute for data curation, and that this effect is scale-dependent.
翻译:通过文档重写生成的合成数据已成为提升语言模型预训练质量的前景技术,然而大多数研究聚焦于英语,且未能系统控制被重写的源数据质量。我们针对葡萄牙语持续预训练场景,系统研究了合成重写与源数据质量之间的交互作用。基于标注了STEM与教育质量得分的葡萄牙语语料库ClassiCC-PT,我们构建了两个不同质量等级的10B词元子集,并使用70亿参数指令微调模型将每个子集重写为四种风格,每种实验条件下生成约400亿词元的合成数据。我们分别在两种条件下训练了以英语为中心的基础模型(11亿与70亿参数),并在涵盖44项任务的综合葡萄牙语基准测试集PoETa V2上进行评估。在70亿参数规模下,重写高质量数据相较于未修改的原始数据获得+3.4个NPM指标的提升,而重写低质量数据仅提升+0.5个NPM。在11亿参数规模下,该交互效应较弱,未修改的低质量数据表现与重写后的高质量数据相当。我们的结果表明:合成重写主要作为质量倍增器而非数据筛选的替代方案,且该效应存在规模依赖性。