In some high-dimensional and semiparametric inference problems involving two populations, the parameter of interest can be characterized by two-sample U-statistics involving some nuisance parameters. In this work we first extend the framework of one-step estimation with cross-fitting to two-sample U-statistics, showing that using an orthogonalized influence function can effectively remove the first order bias, resulting in asymptotically normal estimates of the parameter of interest. As an example, we apply this method and theory to the problem of testing two-sample conditional distributions, also known as strong ignorability. When combined with a conformal-based rank-sum test, we discover that the nuisance parameters can be divided into two categories, where in one category the nuisance estimation accuracy does not affect the testing validity, whereas in the other the nuisance estimation accuracy must satisfy the usual requirement for the test to be valid. We believe these findings provide further insights into and enhance the conformal inference toolbox.
翻译:在一些涉及两个总体的高维和半参数推断问题中,感兴趣的参数可以通过涉及某些干扰参数的双样本U统计量来刻画。本文首先将带交叉拟合的一步估计框架推广至双样本U统计量,表明使用正交化影响函数可以有效消除一阶偏差,从而得到感兴趣参数的渐近正态估计。作为示例,我们将该方法及理论应用于双样本条件分布检验问题(亦称强可忽略性)。当与基于保序的秩和检验相结合时,我们发现干扰参数可分为两类:一类中干扰估计精度不影响检验有效性,而另一类中干扰估计精度必须满足常规要求才能保证检验有效。我们相信这些发现为保序推断工具箱提供了新的见解并对其进行了扩展。