Nonparametric two-sample tests such as the Maximum Mean Discrepancy (MMD) are often used to detect differences between two distributions in machine learning applications. However, the majority of existing literature assumes that error-free samples from the two distributions of interest are available.We relax this assumption and study the estimation of the MMD under $\epsilon$-contamination, where a possibly non-random $\epsilon$ proportion of one distribution is erroneously grouped with the other. We show that under $\epsilon$-contamination, the typical estimate of the MMD is unreliable. Instead, we study partial identification of the MMD, and characterize sharp upper and lower bounds that contain the true, unknown MMD. We propose a method to estimate these bounds, and show that it gives estimates that converge to the sharpest possible bounds on the MMD as sample size increases, with a convergence rate that is faster than alternative approaches. Using three datasets, we empirically validate that our approach is superior to the alternatives: it gives tight bounds with a low false coverage rate.
翻译:非参数双样本检验(如最大均值差异MMD)常被用于机器学习中检测两个分布间的差异。然而,现有文献大多假设可获取感兴趣两个分布的无误差样本。我们放宽了这一假设,研究了在ε-污染条件下的MMD估计问题——其中非随机比例的ε样本从一个分布被错误归入另一个分布。我们证明,在ε-污染条件下,传统的MMD估计结果不可靠。为此,我们转而研究MMD的部分识别,刻画包含真实未知MMD的严苛上下界。我们提出一种估计这些边界的方法,并证明该方法随样本量增大会渐近收敛至MMD的最紧边界,且收敛速度优于其他方法。基于三个数据集的实证验证表明,我们的方法优于现有替代方案:在保持低错误覆盖率的同时,能提供紧致的边界估计。