We propose a new density estimation algorithm. Given $n$ i.i.d. samples from a distribution belonging to a class of densities on $\mathbb{R}^d$, our estimator outputs any density in the class whose ''perceptron discrepancy'' with the empirical distribution is at most $O(\sqrt{d/n})$. The perceptron discrepancy between two distributions is defined as the largest difference in mass that they place on any halfspace of $\mathbb{R}^d$. It is shown that this estimator achieves expected total variation distance to the truth that is almost minimax optimal over the class of densities with bounded Sobolev norm and Gaussian mixtures. This suggests that regularity of the prior distribution could be an explanation for the efficiency of the ubiquitous step in machine learning that replaces optimization over large function spaces with simpler parametric classes (e.g. in the discriminators of GANs). We generalize the above to show that replacing the ''perceptron discrepancy'' with the generalized energy distance of Sz\'ekeley-Rizzo further improves total variation loss. The generalized energy distance between empirical distributions is easily computable and differentiable, thus making it especially useful for fitting generative models. To the best of our knowledge, it is the first example of a distance with such properties for which there are minimax statistical guarantees.
翻译:我们提出了一种新的密度估计算法。给定来自 $\mathbb{R}^d$ 上一类密度分布中独立同分布抽样的 $n$ 个样本,我们的估计器输出该类中任意一个密度,其与经验分布的"感知机差异"至多为 $O(\sqrt{d/n})$。两个分布之间的感知机差异定义为它们分配给 $\mathbb{R}^d$ 中任意半空间的最大质量差。研究表明,该估计器在总变差距离上的期望值几乎达到了有界 Sobolev 范数和高斯混合密度类上的极小极大最优。这表明先验分布的正则性可能是机器学习中常用步骤(如生成对抗网络的判别器)效率的解释,该步骤将大函数空间上的优化替换为更简单的参数类。我们将上述内容推广,证明将"感知机差异"替换为 Székely-Rizzo 的广义能量距离能进一步降低总变差损失。经验分布间的广义能量距离易于计算且可微,因此特别适用于拟合生成模型。据我们所知,这是首个在具有极小极大统计保证条件下同时具备这些特性的距离度量。