We use a suitable version of the so-called "kernel trick" to devise two-sample (homogeneity) tests, especially focussed on high-dimensional and functional data. Our proposal entails a simplification related to the important practical problem of selecting an appropriate kernel function. Specifically, we apply a uniform variant of the kernel trick which involves the supremum within a class of kernel-based distances. We obtain the asymptotic distribution (under the null and alternative hypotheses) of the test statistic. The proofs rely on empirical processes theory, combined with the delta method and Hadamard (directional) differentiability techniques, and functional Karhunen-Lo\`eve-type expansions of the underlying processes. This methodology has some advantages over other standard approaches in the literature. We also give some experimental insight into the performance of our proposal compared to the original kernel-based approach \cite{Gretton2007} and the test based on energy distances \cite{Szekely-Rizzo-2017}.
翻译:我们采用适当版本的所谓“核技巧”来设计两样本(同质性)检验,特别针对高维和函数型数据。我们的方案简化了与选择合适核函数这一重要实际问题相关的步骤。具体而言,我们应用核技巧的均匀变体,该变体涉及一类基于核的距离的 supremum。我们得到了检验统计量的渐近分布(在原假设和备择假设下)。证明依赖于经验过程理论,结合 delta 方法和 Hadamard(方向)可微性技术,以及底层过程的函数型 Karhunen-Loève 展开。该方法相比文献中的其他标准方法具有一定优势。我们还通过与原始基于核的方法 \cite{Gretton2007} 和基于能量距离的检验 \cite{Szekely-Rizzo-2017} 进行比较,给出了关于我们方案性能的实验性见解。