We introduce kernel thinning, a new procedure for compressing a distribution $\mathbb{P}$ more effectively than i.i.d. sampling or standard thinning. Given a suitable reproducing kernel $\mathbf{k}_{\star}$ and $\mathcal{O}(n^2)$ time, kernel thinning compresses an $n$-point approximation to $\mathbb{P}$ into a $\sqrt{n}$-point approximation with comparable worst-case integration error across the associated reproducing kernel Hilbert space. The maximum discrepancy in integration error is $\mathcal{O}_d(n^{-1/2}\sqrt{\log n})$ in probability for compactly supported $\mathbb{P}$ and $\mathcal{O}_d(n^{-\frac{1}{2}} (\log n)^{(d+1)/2}\sqrt{\log\log n})$ for sub-exponential $\mathbb{P}$ on $\mathbb{R}^d$. In contrast, an equal-sized i.i.d. sample from $\mathbb{P}$ suffers $\Omega(n^{-1/4})$ integration error. Our sub-exponential guarantees resemble the classical quasi-Monte Carlo error rates for uniform $\mathbb{P}$ on $[0,1]^d$ but apply to general distributions on $\mathbb{R}^d$ and a wide range of common kernels. Moreover, the same construction delivers near-optimal $L^\infty$ coresets in $\mathcal O(n^2)$ time. We use our results to derive explicit non-asymptotic maximum mean discrepancy bounds for Gaussian, Mat\'ern, and B-spline kernels and present two vignettes illustrating the practical benefits of kernel thinning over i.i.d. sampling and standard Markov chain Monte Carlo thinning, in dimensions $d=2$ through $100$.
翻译:我们提出核稀疏化(kernel thinning)这一新方法,用于比独立同分布采样或标准稀疏化更有效地压缩分布$\mathbb{P}$。给定合适的再生核$\mathbf{k}_{\star}$及$\mathcal{O}(n^2)$时间,核稀疏化能将$\mathbb{P}$的$n$点近似压缩为$\sqrt{n}$点近似,且在关联再生核希尔伯特空间上保持具有可比性的最坏情况积分误差。对于紧支撑分布$\mathbb{P}$,积分误差的最大差异概率达到$\mathcal{O}_d(n^{-1/2}\sqrt{\log n})$;对于$\mathbb{R}^d$上的次指数分布$\mathbb{P}$,该误差为$\mathcal{O}_d(n^{-\frac{1}{2}} (\log n)^{(d+1)/2}\sqrt{\log\log n})$。相比之下,等大小的$\mathbb{P}$独立同分布样本的积分误差为$\Omega(n^{-1/4})$。我们的次指数保证类似于$[0,1]^d$上均匀分布$\mathbb{P}$的经典拟蒙特卡洛误差率,但适用于$\mathbb{R}^d$上的一般分布及多种常见核。此外,相同构造可在$\mathcal O(n^2)$时间内获得近最优$L^\infty$核心集。我们利用这些结果推导了高斯核、马特恩核和B样条核的显式非渐近最大均值差异界,并通过两个案例展示了在维度$d=2$至$100$中,核稀疏化相比独立同分布采样和标准马尔可夫链蒙特卡洛稀疏化的实际优势。