Markov chain Monte Carlo (MCMC) is a class of general-purpose algorithms for sampling from unnormalized densities. There are two well-known problems facing MCMC in high dimensions: (i) The distributions of interest are concentrated in pockets separated by large regions with small probability mass, and (ii) The log-concave pockets themselves are typically ill-conditioned. We introduce a framework to tackle these problems using isotropic Gaussian smoothing. We prove one can always decompose sampling from a density (minimal assumptions made on the density) into a sequence of sampling from log-concave conditional densities via accumulation of noisy measurements with equal noise levels. This construction keeps track of a history of samples, making it non-Markovian as a whole, but the history only shows up in the form of an empirical mean, making the memory footprint minimal. Our sampling algorithm generalizes walk-jump sampling [1]. The "walk" phase becomes a (non-Markovian) chain of log-concave Langevin chains. The "jump" from the accumulated measurements is obtained by empirical Bayes. We study our sampling algorithm quantitatively using the 2-Wasserstein metric and compare it with various Langevin MCMC algorithms. We also report a remarkable capacity of our algorithm to "tunnel" between modes of a distribution.
翻译:马尔可夫链蒙特卡洛(MCMC)是一类从未归一化密度中采样的通用算法。高维MCMC面临两个众所周知的问题:(i)目标分布集中在由大区域低概率质量分隔的"口袋"中;(ii)对数凹口袋本身通常呈现病态条件。我们提出一个利用各向同性高斯平滑处理这些问题的框架。我们证明,通过等噪声水平的噪声测量积累,总可以将密度采样(对密度施加最少假设)分解为一系列从对数凹条件密度采样的过程。该结构通过追踪样本历史记录形成整体非马尔可夫性,但历史仅以经验均值形式存在,使得内存占用极小。我们的采样算法推广了跳跃采样[1]:"行走"阶段成为对数凹朗之万链的(非马尔可夫)链,而基于积累测量结果的"跳跃"通过经验贝叶斯实现。我们使用2-瓦瑟斯坦度量对采样算法进行定量分析,并与多种朗之万MCMC算法进行比较。我们还报告了该算法在分布模态间实现"隧穿"的显著能力。