We advocate for a new paradigm of cosmological likelihood-based inference, leveraging recent developments in machine learning and its underlying technology, to accelerate Bayesian inference in high-dimensional settings. Specifically, we combine (i) emulation, where a machine learning model is trained to mimic cosmological observables, e.g. CosmoPower-JAX; (ii) differentiable and probabilistic programming, e.g. JAX and NumPyro, respectively; (iii) scalable Markov chain Monte Carlo (MCMC) sampling techniques that exploit gradients, e.g. Hamiltonian Monte Carlo; and (iv) decoupled and scalable Bayesian model selection techniques that compute the Bayesian evidence purely from posterior samples, e.g. the learned harmonic mean implemented in harmonic. This paradigm allows us to carry out a complete Bayesian analysis, including both parameter estimation and model selection, in a fraction of the time of traditional approaches. First, we demonstrate the application of this paradigm on a simulated cosmic shear analysis for a Stage IV survey in 37- and 39-dimensional parameter spaces, comparing $\Lambda$CDM and a dynamical dark energy model ($w_0w_a$CDM). We recover posterior contours and evidence estimates that are in excellent agreement with those computed by the traditional nested sampling approach while reducing the computational cost from 8 months on 48 CPU cores to 2 days on 12 GPUs. Second, we consider a joint analysis between three simulated next-generation surveys, each performing a 3x2pt analysis, resulting in 157- and 159-dimensional parameter spaces. Standard nested sampling techniques are simply not feasible in this high-dimensional setting, requiring a projected 12 years of compute time on 48 CPU cores; on the other hand, the proposed approach only requires 8 days of compute time on 24 GPUs. All packages used in our analyses are publicly available.
翻译:我们倡导一种基于似然的宇宙学推断新范式,通过利用机器学习及其底层技术的最新进展,加速高维贝叶斯推断。具体而言,该方法结合了以下技术:(i)模拟仿真(emulation),即训练机器学习模型以模仿宇宙学可观测量,例如CosmoPower-JAX;(ii)可微与概率编程,分别基于JAX和NumPyro框架;(iii)利用梯度信息的可扩展马尔可夫链蒙特卡洛(MCMC)采样技术,例如哈密顿蒙特卡洛方法;以及(iv)仅从后验样本计算贝叶斯证据的解耦式可扩展贝叶斯模型选择技术,例如通过harmonic实现的学习调和均值方法。该范式使我们能够在传统方法所需时间的一小部分内,完成包括参数估计和模型选择在内的完整贝叶斯分析。首先,我们通过模拟的弱引力透镜巡天分析展示了该范式的应用:针对第四阶段巡天项目,在37维和39维参数空间中比较了$\Lambda$CDM模型与动态暗能量模型($w_0w_a$CDM)。结果显示,后验轮廓和证据估计值与传统嵌套采样方法的结果高度一致,同时计算成本从48个CPU核心运行8个月降至12个GPU运行2天。其次,我们考虑了三项模拟的下一代巡天联合分析(每项均执行3×2点分析),形成157维和159维参数空间。标准嵌套采样技术在此高维场景中完全不可行——预计需48个CPU核心运行12年;而本方法仅需24个GPU运行8天。本研究中使用的所有软件包均已公开。