Generative artificial intelligence is rapidly transforming the supply side of training data: an increasing share of new tokens, images, and structured records is produced by previous-generation models rather than by human originators. Recursive training on such synthetic content induces a measurable and often irreversible loss of distributional fidelity, a phenomenon known as model collapse. We develop the first unified microeconomic theory of synthetic data markets under model collapse. We introduce the Synthetic Data Contamination Equilibrium (SDCE), prove existence and generic uniqueness, derive a welfare decomposition W = W_prod + W_cons - L_coll - L_info, establish a Wasserstein-gradient-flow mean-field collapse limit, prove an impossibility of information-constrained implementation, and obtain closed-form expressions for the welfare-maximizing provenance subsidy s* = KL(q||p)/(2 kappa) and the welfare-maximizing watermark strength w* = (1 - psi) KL(q||p)/(2 kappa psi). We prove an information-theoretic Cramer-Rao lower bound on any provenance estimator using only producer-side observations and show that the Provenance-Market Iterative Retraining (PMIR) algorithm attains this bound up to constants while converging to an epsilon-SDCE in O(epsilon^-2 log T) iterations. A reduced-form OLS estimation on a C4-synthetic benchmark over ten retraining generations yields a collapse-rate coefficient b-hat = 0.181 (HAC s.e. 0.024), within one standard error of the structural prediction 0.183. Calibrated experiments raise generation-ten model quality by 23.1 percent over the unregulated benchmark while lowering the 2-Wasserstein drift on a held-out diversity probe from 0.318 to 0.142. Scaling experiments over generations t in {1,...,10} recover a logarithmic-in-t collapse law log Q_t = log Q_0 - 0.183 t rho^2 with R^2 = 0.962.
翻译:生成式人工智能正迅速改变训练数据的供给侧:越来越多的新词元、图像和结构化记录由前代模型而非人类原始创作者产生。对此类合成内容进行递归训练会导致分布保真度出现可度量且往往不可逆的损失,这一现象被称为模型坍缩。我们首次建立了模型坍缩下合成数据市场的统一微观经济学理论。我们提出合成数据污染均衡(SDCE),证明其存在性与一般唯一性,推导出福利分解式 W = W_prod + W_cons - L_coll - L_info,建立了Wasserstein梯度流平均场坍缩极限,证明了信息约束下实施的不可能性,并得到了福利最大化溯源补贴 s* = KL(q||p)/(2 κ) 与福利最大化水印强度 w* = (1 - ψ) KL(q||p)/(2 κ ψ) 的闭式解。我们证明了仅利用生产者侧观测值的任何溯源估计量满足信息论Cramér-Rao下界,并证明溯源市场迭代再训练(PMIR)算法在O(ε⁻² log T)次迭代内收敛至ε-SDCE的同时,以常数倍率逼近该下界。基于C4合成基准数据集在十次再训练代际上的简约形式OLS估计,得到坍缩速率系数 ˆβ = 0.181(HAC标准误0.024),与结构参数预测值0.183相差在一个标准误内。校准实验将第十代模型质量较无监管基准提升23.1%,同时将保留多样性探针上的2-Wasserstein漂移从0.318降至0.142。在t∈{1,...,10}代际上的标度实验恢复了对数时间坍缩律 log Q_t = log Q_0 - 0.183 t ρ²,R² = 0.962。