Data assimilation, in its most comprehensive form, addresses the Bayesian inverse problem of identifying plausible state trajectories that explain noisy or incomplete observations of stochastic dynamical systems. Various approaches have been proposed to solve this problem, including particle-based and variational methods. However, most algorithms depend on the transition dynamics for inference, which becomes intractable for long time horizons or for high-dimensional systems with complex dynamics, such as oceans or atmospheres. In this work, we introduce score-based data assimilation for trajectory inference. We learn a score-based generative model of state trajectories based on the key insight that the score of an arbitrarily long trajectory can be decomposed into a series of scores over short segments. After training, inference is carried out using the score model, in a non-autoregressive manner by generating all states simultaneously. Quite distinctively, we decouple the observation model from the training procedure and use it only at inference to guide the generative process, which enables a wide range of zero-shot observation scenarios. We present theoretical and empirical evidence supporting the effectiveness of our method.
翻译:数据同化在其最完整的形式中,解决的是贝叶斯逆问题,即识别能够解释随机动力系统中含有噪声或不完全观测的可行状态轨迹。已有多种方法被提出以解决这一问题,包括基于粒子和变分方法。然而,大多数算法依赖于转移动力学进行推断,这在长时间范围或具有复杂动力学的高维系统(如海洋或大气)中变得不可处理。在本工作中,我们引入基于分数的数据同化用于轨迹推断。基于一个关键洞见——任意长时间轨迹的分数可以分解为一系列短段上的分数——我们学习一个基于分数的状态轨迹生成模型。训练后,通过同时生成所有状态的非自回归方式,利用分数模型进行推断。独特的是,我们将观测模型与训练过程解耦,并仅在推断时用于引导生成过程,从而实现广泛的零样本观测场景。我们提供了支持该方法有效性的理论和实证证据。