The pretraining data mixture of Large Language Models (LLMs) constitutes their "digital DNA", shaping model behaviors, capabilities, and failure modes. Yet this composition is rarely disclosed, making post-hoc auditing of data combination or provenance difficult. In this work, we formalize $\textbf{Data Mixture Surgery (DMS)}$: given only generated text from a target LLM, estimate the domain-level distribution of its pretraining corpus under a predefined taxonomy. We propose $\textbf{LLMSurgeon}$, a strong framework that casts DMS as an inverse problem under the label-shift assumption. Rather than directly aggregating classifier outputs, LLMSurgeon estimates a calibrated $\textit{soft}$ confusion matrix and solves a constrained inverse problem to correct systematic domain confusion and recover the latent mixture prior. To evaluate, we introduce $\textbf{LLMScan}$, a recipe-verifiable evaluation suite built from open-source LLMs with transparent pretraining mixtures. Across LLMScan, LLMSurgeon recovers domain mixtures with high fidelity under fixed protocols. Our work presents a practical, post-hoc approach for auditing the digital DNA of foundation models without access to their training data.
翻译:大语言模型(LLM)的预训练数据混合构成了其“数字DNA”,塑造了模型的行为、能力与失效模式。然而,这一组成极少被公开,使得事后审计数据的组合或来源变得困难。在本工作中,我们形式化了$\textbf{数据混合手术(Data Mixture Surgery, DMS)}$:仅根据目标LLM生成的文本,在预定义分类体系下估算其预训练语料的领域级分布。我们提出$\textbf{LLMSurgeon}$,一个将DMS建模为标签偏移假设下反问题的强健框架。LLMSurgeon不直接聚合分类器输出,而是估计校准后的$\textit{软}$混淆矩阵,并求解带约束的反问题以修正系统性领域混淆并恢复潜在的混合先验。为进行评估,我们引入$\textbf{LLMScan}$,一个基于预训练混合透明的开源LLM构建的配方可验证评估套件。在LLMScan上,LLMSurgeon在固定协议下以高保真度恢复领域混合。我们的工作提供了一种无需访问训练数据即可事后审计基础模型数字DNA的实用方法。