Many problems in computational science and engineering become one-to-many after coarse graining, partial observation, or inverse reconstruction: a resolved state may not determine a unique subgrid forcing, a structural descriptor may not determine a unique effective response, and a low-resolution observation may correspond to many plausible high-resolution fields. In such settings, deterministic surrogates may learn a well-defined mathematical object while still missing application-relevant uncertainty. This tutorial develops a self-contained module centered on the conditional-mean barrier: the point at which a squared-loss predictor has reached the conditional mean and the remaining error is irreducible aleatoric variance. We give two diagnostics for locating this barrier, residual-feature orthogonality and the coefficient of determination against its explained-variance ceiling, and prove that adding latent randomness to a squared-loss predictor collapses it back to the conditional mean. Crossing the barrier therefore requires a loss that scores distributions rather than point predictions. We briefly organize common distributional objectives, including negative log-likelihood, moment and observable matching, variational objectives, adversarial divergences, and score matching, by the feature of the conditional law each targets. The emphasis is the boundary itself and a finite-data procedure for recognizing it, rather than a survey of methods beyond it. CPU-based demonstrations on a two-branch law and a two-scale Lorenz-96 closure problem show how the diagnostics distinguish deterministic underfitting from residual distributional variability.


翻译:计算科学与工程中的许多问题在粗粒化、部分观测或逆重建后会变成一对多问题:一个解析状态可能无法确定唯一的子网格强迫项,一个结构描述符可能无法确定唯一的有效响应,而一个低分辨率观测可能对应多个合理的高分辨率场。在此类场景中,确定性代理可能学习到一个定义良好的数学对象,但仍会遗漏与应用相关的不确定性。本教程开发了一个自包含的模块,聚焦于条件均值障碍:即平方损失预测器达到条件均值,且剩余误差为不可约的偶然性方差时的临界点。我们提出了两种定位该障碍的诊断方法——残差特征正交性与针对可解释方差上限的拟合优度系数,并证明了向平方损失预测器中添加潜在随机性会使其坍塌回条件均值。因此,跨越此障碍需要一种对分布而非点预测进行评分的损失函数。我们简要整理了常见的分布目标,包括负对数似然、矩与可观测量匹配、变分目标、对抗散度以及分数匹配,根据每种方法所针对的条件分布特征进行分类。重点在于障碍本身及识别它的有限数据程序,而非对超越它的方法进行综述。基于CPU的双分支定律与双尺度Lorenz-96闭合问题演示展示了诊断如何区分确定性欠拟合与残差分布变异性。

0
下载
关闭预览

相关内容

【牛津大学博士论文】机器学习中的对称性与泛化
专知会员服务
22+阅读 · 2025年1月8日
「PPT」深度学习中的不确定性估计
专知
27+阅读 · 2019年7月20日
机器学习中如何处理不平衡数据?
机器之心
13+阅读 · 2019年2月17日
人工智能顶刊TPAMI2019最新《多模态机器学习综述》
人工智能学家
29+阅读 · 2019年1月19日
从信息瓶颈理论一瞥机器学习的“大一统理论”
推荐|机器学习中的模型评价、模型选择和算法选择!
全球人工智能
10+阅读 · 2018年2月5日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
《最强大的军事网状网络》
专知会员服务
3+阅读 · 9月7日
《预测陆军征兵任务分配》110页
专知会员服务
4+阅读 · 9月7日
分层反无人机系统发展新趋势
专知会员服务
11+阅读 · 9月3日
何为协作武器?
专知会员服务
11+阅读 · 9月1日
相关VIP内容
【牛津大学博士论文】机器学习中的对称性与泛化
专知会员服务
22+阅读 · 2025年1月8日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员