Bayesian inference provides a powerful tool for leveraging observational data to inform model predictions and uncertainties. However, when such data is limited, Bayesian inference may not adequately constrain uncertainty without the use of highly informative priors. Common approaches for constructing informative priors typically rely on either assumptions or knowledge of the underlying physics, which may not be available in all scenarios. In this work, we consider the scenario where data are available on a population of assets/individuals, which occurs in many problem domains such as biomedical or digital twin applications, and leverage this population-level data to systematically constrain the Bayesian prior and subsequently improve individualized inferences. The approach proposed in this paper is based upon a recently developed technique known as data-consistent inversion (DCI) for constructing a pullback probability measure. Succinctly, we utilize DCI to build population-informed priors for subsequent Bayesian inference on individuals. While the approach is general and applies to nonlinear maps and arbitrary priors, we prove that for linear inverse problems with Gaussian priors, the population-informed prior produces an increase in the information gain as measured by the determinant and trace of the inverse posterior covariance. We also demonstrate that the Kullback-Leibler divergence often improves with high probability. Numerical results, including linear-Gaussian examples and one inspired by digital twins for additively manufactured assets, indicate that there is significant value in using these population-informed priors.
翻译:贝叶斯推断为利用观测数据指导模型预测与不确定性量化提供了强大工具。然而当数据有限时,若无高度信息性先验的辅助,贝叶斯推断可能无法充分约束不确定性。构建信息性先验的常用方法通常依赖于对底层物理机制的假设或认知,但这在诸多场景中并不可得。本研究针对资产/个体群体数据可获取的场景(常见于生物医学或数字孪生等应用领域),利用此类群体层级数据系统性地约束贝叶斯先验分布,进而提升个体化推断效果。本文提出的方法基于近期发展的数据一致性反演技术,该技术用于构建拉回概率测度。简言之,我们运用DCI构建群体信息先验,以支持后续针对个体的贝叶斯推断。虽然该方法具有普适性,适用于非线性映射与任意先验分布,但我们在线性反问题与高斯先验的特定条件下证明:群体信息先验能通过后验协方差逆矩阵的行列式与迹度量的信息增益提升推断效果。同时我们证明Kullback-Leibler散度在大概率下亦会改善。数值实验结果(包括线性高斯示例及受增材制造资产数字孪生启发的案例)表明,使用此类群体信息先验具有显著价值。