Kernel two-sample tests have been widely used for multivariate data in testing equal distribution. However, existing tests based on mapping distributions into a reproducing kernel Hilbert space are mainly targeted at specific alternatives and do not work well for some scenarios when the dimension of the data is moderate to high due to the curse of dimensionality. We propose a new test statistic that makes use of a common pattern under moderate and high dimensions and achieves substantial power improvements over existing kernel two-sample tests for a wide range of alternatives. We also propose alternative testing procedures that maintain high power with low computational cost, offering easy off-the-shelf tools for large datasets. The new approaches are compared to other state-of-the-art tests under various settings and show good performance. The new approaches are illustrated on two applications: The comparison of musks and non-musks using the shape of molecules, and the comparison of taxi trips started from John F.Kennedy airport in consecutive months. All proposed methods are implemented in an R package kerTests.
翻译:核双样本检验已广泛应用于多元数据中的等分布检验。然而,现有基于将分布映射至再生核希尔伯特空间的方法主要针对特定备择假设设计,且由于维度灾难,当数据维度处于中等或较高水平时,其在某些场景下效果不佳。我们提出一种利用中高维度常见模式的新检验统计量,能在广泛备择假设下显著提升现有核双样本检验的统计功效。同时,我们提出替代性检验流程,在保持高功效的同时降低计算成本,为大型数据集提供便捷的即用型工具。新方法在多种设置下与其它最新检验方法进行比较,展现出优异性能。我们通过两个应用案例展示新方法:利用分子形状比较麝香类与非麝香类物质,以及比较约翰·F·肯尼迪机场连续月份出发的出租车行程。所有提出方法均已集成于R包kerTests中。