Stochastic gradients have been widely integrated into Langevin-based methods to improve their scalability and efficiency in solving large-scale sampling problems. However, the proximal sampler, which exhibits much faster convergence than Langevin-based algorithms in the deterministic setting Lee et al. (2021), has yet to be explored in its stochastic variants. In this paper, we study the Stochastic Proximal Samplers (SPS) for sampling from non-log-concave distributions. We first establish a general framework for implementing stochastic proximal samplers and establish the convergence theory accordingly. We show that the convergence to the target distribution can be guaranteed as long as the second moment of the algorithm trajectory is bounded and restricted Gaussian oracles can be well approximated. We then provide two implementable variants based on Stochastic gradient Langevin dynamics (SGLD) and Metropolis-adjusted Langevin algorithm (MALA), giving rise to SPS-SGLD and SPS-MALA. We further show that SPS-SGLD and SPS-MALA can achieve $\epsilon$-sampling error in total variation (TV) distance within $\tilde{\mathcal{O}}(d\epsilon^{-2})$ and $\tilde{\mathcal{O}}(d^{1/2}\epsilon^{-2})$ gradient complexities, which outperform the best-known result by at least an $\tilde{\mathcal{O}}(d^{1/3})$ factor. This enhancement in performance is corroborated by our empirical studies on synthetic data with various dimensions, demonstrating the efficiency of our proposed algorithm.
翻译:随机梯度已被广泛集成到基于朗之万的方法中,以提高其在解决大规模采样问题时的可扩展性和效率。然而,近端采样器——在确定性设置下表现出比基于朗之万的算法快得多的收敛速度(Lee等人,2021)——其随机变体尚未得到探索。本文研究了用于从非对数凹分布采样的随机近端采样器(SPS)。我们首先建立了一个实现随机近端采样器的通用框架,并相应地建立了收敛理论。我们证明,只要算法轨迹的二阶矩有界且受限高斯预言机能够得到良好近似,就能保证收敛到目标分布。随后,我们基于随机梯度朗之万动力学(SGLD)和Metropolis调整朗之万算法(MALA)提供了两种可实现的变体,即SPS-SGLD和SPS-MALA。我们进一步证明,SPS-SGLD和SPS-MALA在总变差(TV)距离上达到$\epsilon$采样误差所需的梯度复杂度分别为$\tilde{\mathcal{O}}(d\epsilon^{-2})$和$\tilde{\mathcal{O}}(d^{1/2}\epsilon^{-2})$,这比目前已知的最佳结果至少提高了$\tilde{\mathcal{O}}(d^{1/3})$倍。我们在不同维度的合成数据上进行的实证研究证实了这种性能提升,证明了所提算法的高效性。