Transformer-based scientific foundation models are increasingly deployed in high-stakes settings, but current architectures give deterministic outputs and provide limited support for calibrated predictive uncertainty. We propose Stochastic Attention, a sample average lightweight inference-time modification that randomizes attention by replacing softmax weights with normalized multinomial samples controlled by a single concentration parameter, and produces predictive ensembles without retraining. To set this parameter, we introduce a calibration objective that matches the stochastic attention output with the target, yielding an efficient univariate post-hoc tuning problem. We evaluate this mechanism on scientific foundation models for weather and time-series forecasting, as well as several regression tasks. Across benchmarks against uncertainty-aware baselines, we find that Sample Average Stochastic Attention achieves the strongest native calibration and the sharpest prediction intervals at comparable calibration, with adaptation costs nearly three orders of magnitude lower than the next-best baseline.
翻译:基于Transformer的科学基础模型正越来越多地应用于高风险场景,但现有架构输出确定性结果,且对校准后的预测不确定性支持有限。我们提出随机注意力机制——一种基于样本均值的轻量级推理时改进方法:通过用受单一浓度参数控制的归一化多项式样本替代Softmax权重来随机化注意力,无需重新训练即可生成预测集成。为设定该参数,我们引入校准目标函数使随机注意力输出与目标值匹配,从而形成高效的单变量后验调参问题。我们在天气与时间序列预测,以及多项回归任务的科学基础模型上评估该机制。通过与不确定性感知基线方法的基准对比发现:在相当校准水平下,样本平均随机注意力实现了最强的原生校准性和最窄的预测区间,其自适应成本比次优基线低近三个数量级。