As the demand to integrate Artificial Intelligence into high-stakes environments continues to grow, explaining the reasoning behind neural-network predictions has shifted from a theoretical curiosity to a strict operational requirement. Our work is motivated by the explanations of autoregressive neural predictions on dynamic physical fields, as in weather forecasting. Gradient-based feature attribution methods are widely used to explain the predictions on such data, in particular due to their scalability to high-dimensional inputs. It is also interesting to remark that gradient-based techniques such as SmoothGrad are now standard on images to robustify the explanations using pointwise averages of the attribution maps obtained from several noised inputs. Our goal is to efficiently adapt this aggregation strategy to dynamic physical fields. To do so, our first contribution is to identify a fundamental failure mode when averaging perturbed attribution maps on dynamic physical fields: stochastic input perturbations do not induce stationary amplitude noise in attribution maps, but instead cause a geometric displacement of the attributions. Consequently, pointwise averaging blurs these spatially misaligned features. To tackle this issue, we introduce WassersteinGrad, which extracts a geometric consensus of perturbed attribution maps by computing their entropic Wasserstein barycenter. The results, obtained on regional weather data and a meteorologist-validated neural model, demonstrate promising explainability properties of WassersteinGrad over gradient-based baselines across both single-step and autoregressive forecasting settings.
翻译:随着将人工智能融入高风险环境的需求持续增长,解释神经网络预测背后的推理已从理论探索转变为严格的运营要求。我们的工作源于对动态物理场自回归神经网络预测的解释,例如天气预报。梯度特征归因方法被广泛用于解释此类数据的预测,特别是因其对高维输入的可扩展性。值得注意的是,诸如SmoothGrad等梯度技术已成为图像领域的标准方法,通过对多个噪声输入获得的归因图进行逐点平均来增强解释的鲁棒性。我们的目标是高效地将这种聚合策略适配到动态物理场上。为此,我们的首要贡献是识别出在动态物理场上平均扰动归因图时的基本失效模式:随机输入扰动不会在归因图中产生平稳的幅度噪声,而是导致归因的几何位移。因此,逐点平均会模糊这些空间错位的特征。为应对这一问题,我们引入WassersteinGrad方法,通过计算扰动归因图的熵正则化Wasserstein重心来提取其几何共识。在区域天气数据及经气象学家验证的神经网络模型上获得的结果表明,在单步预测和自回归预测场景中,WassersteinGrad相较于基于梯度的基线方法展现出更优的可解释性。