Generative artificial intelligence is rapidly transforming the supply side of training data: an increasing share of new tokens, images, and structured records is produced by previous-generation models rather than by human originators. Recursive training on such synthetic content induces a measurable and often irreversible loss of distributional fidelity, a phenomenon known as model collapse. We develop the first unified microeconomic theory of synthetic data markets under model collapse. We introduce the Synthetic Data Contamination Equilibrium (SDCE), prove existence and generic uniqueness, derive a welfare decomposition W = W_prod + W_cons - L_coll - L_info, establish a Wasserstein-gradient-flow mean-field collapse limit, prove an impossibility of information-constrained implementation, and obtain closed-form expressions for the welfare-maximizing provenance subsidy s* = KL(q||p)/(2 kappa) and the welfare-maximizing watermark strength w* = (1 - psi) KL(q||p)/(2 kappa psi). We prove an information-theoretic Cramer-Rao lower bound on any provenance estimator using only producer-side observations and show that the Provenance-Market Iterative Retraining (PMIR) algorithm attains this bound up to constants while converging to an epsilon-SDCE in O(epsilon^-2 log T) iterations. A reduced-form OLS estimation on a C4-synthetic benchmark over ten retraining generations yields a collapse-rate coefficient b-hat = 0.181 (HAC s.e. 0.024), within one standard error of the structural prediction 0.183. Calibrated experiments raise generation-ten model quality by 23.1 percent over the unregulated benchmark while lowering the 2-Wasserstein drift on a held-out diversity probe from 0.318 to 0.142. Scaling experiments over generations t in {1,...,10} recover a logarithmic-in-t collapse law log Q_t = log Q_0 - 0.183 t rho^2 with R^2 = 0.962.


翻译:生成式人工智能正迅速改变训练数据的供给侧:越来越多的新词元、图像和结构化记录由前代模型而非人类原始创作者产生。对此类合成内容进行递归训练会导致分布保真度出现可度量且往往不可逆的损失,这一现象被称为模型坍缩。我们首次建立了模型坍缩下合成数据市场的统一微观经济学理论。我们提出合成数据污染均衡(SDCE),证明其存在性与一般唯一性,推导出福利分解式 W = W_prod + W_cons - L_coll - L_info,建立了Wasserstein梯度流平均场坍缩极限,证明了信息约束下实施的不可能性,并得到了福利最大化溯源补贴 s* = KL(q||p)/(2 κ) 与福利最大化水印强度 w* = (1 - ψ) KL(q||p)/(2 κ ψ) 的闭式解。我们证明了仅利用生产者侧观测值的任何溯源估计量满足信息论Cramér-Rao下界,并证明溯源市场迭代再训练(PMIR)算法在O(ε⁻² log T)次迭代内收敛至ε-SDCE的同时,以常数倍率逼近该下界。基于C4合成基准数据集在十次再训练代际上的简约形式OLS估计,得到坍缩速率系数 ˆβ = 0.181(HAC标准误0.024),与结构参数预测值0.183相差在一个标准误内。校准实验将第十代模型质量较无监管基准提升23.1%,同时将保留多样性探针上的2-Wasserstein漂移从0.318降至0.142。在t∈{1,...,10}代际上的标度实验恢复了对数时间坍缩律 log Q_t = log Q_0 - 0.183 t ρ²,R² = 0.962。

0
下载
关闭预览

相关内容

ACM/IEEE第23届模型驱动工程语言和系统国际会议,是模型驱动软件和系统工程的首要会议系列,由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来,模型涵盖了建模的各个方面,从语言和方法到工具和应用程序。模特的参加者来自不同的背景,包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛,参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会,并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。 官网链接:http://www.modelsconference.org/
生成模型中持续学习的综合综述
专知会员服务
25+阅读 · 2025年6月17日
《大语言模型的数据合成与增强综述》
专知会员服务
44+阅读 · 2024年10月19日
《生成式人工智能模型:机遇与风险》
专知会员服务
81+阅读 · 2024年4月22日
谷歌最新《大语言模型合成数据的最佳实践和经验教训》
经济学中的数据科学,Data Science in Economics,附22页pdf
专知会员服务
36+阅读 · 2020年4月1日
模型压缩 | 知识蒸馏经典解读
AINLP
11+阅读 · 2020年5月31日
斯坦福CS236-深度生成模型2019-全套课程资料分享
深度学习与NLP
20+阅读 · 2019年8月20日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
7+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
14+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
13+阅读 · 8月8日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员