Many problems in computational science and engineering become one-to-many after coarse graining, partial observation, or inverse reconstruction: a resolved state may not determine a unique subgrid forcing, a structural descriptor may not determine a unique effective response, and a low-resolution observation may correspond to many plausible high-resolution fields. In such settings, deterministic surrogates may learn a well-defined mathematical object while still missing application-relevant uncertainty. This tutorial develops a self-contained module centered on the conditional-mean barrier: the point at which a squared-loss predictor has reached the conditional mean and the remaining error is irreducible aleatoric variance. We give two diagnostics for locating this barrier, residual-feature orthogonality and the coefficient of determination against its explained-variance ceiling, and prove that adding latent randomness to a squared-loss predictor collapses it back to the conditional mean. Crossing the barrier therefore requires a loss that scores distributions rather than point predictions. We briefly organize common distributional objectives, including negative log-likelihood, moment and observable matching, variational objectives, adversarial divergences, and score matching, by the feature of the conditional law each targets. The emphasis is the boundary itself and a finite-data procedure for recognizing it, rather than a survey of methods beyond it. CPU-based demonstrations on a two-branch law and a two-scale Lorenz-96 closure problem show how the diagnostics distinguish deterministic underfitting from residual distributional variability.
翻译:计算科学与工程中的许多问题在粗粒化、部分观测或逆重建后会变成一对多问题:一个解析状态可能无法确定唯一的子网格强迫项,一个结构描述符可能无法确定唯一的有效响应,而一个低分辨率观测可能对应多个合理的高分辨率场。在此类场景中,确定性代理可能学习到一个定义良好的数学对象,但仍会遗漏与应用相关的不确定性。本教程开发了一个自包含的模块,聚焦于条件均值障碍:即平方损失预测器达到条件均值,且剩余误差为不可约的偶然性方差时的临界点。我们提出了两种定位该障碍的诊断方法——残差特征正交性与针对可解释方差上限的拟合优度系数,并证明了向平方损失预测器中添加潜在随机性会使其坍塌回条件均值。因此,跨越此障碍需要一种对分布而非点预测进行评分的损失函数。我们简要整理了常见的分布目标,包括负对数似然、矩与可观测量匹配、变分目标、对抗散度以及分数匹配,根据每种方法所针对的条件分布特征进行分类。重点在于障碍本身及识别它的有限数据程序,而非对超越它的方法进行综述。基于CPU的双分支定律与双尺度Lorenz-96闭合问题演示展示了诊断如何区分确定性欠拟合与残差分布变异性。