We study the problem of testing whether two conditional distributions are equal using generative models. The proposed method learns a conditional generator from each sample and uses it to create responses at covariate values observed in the other sample, allowing generated and observed responses to be compared directly. By aligning covariates through cross-generation, the approach avoids conditional density-ratio estimation and local smoothing over high-dimensional covariates. The population version of this construction yields a conditional discrepancy that characterizes equality of the two conditional distributions under suitable overlap conditions, while the sample version leads to a test statistic defined as the supremum of an RKHS-indexed empirical process with multiplier bootstrap calibration. A computationally efficient algorithm for evaluating the statistic and its bootstrap analogue is developed based on alternating maximization and the kernel trick. Theoretically, we derive the limiting distribution of the test statistic under both the null and alternative hypotheses, prove bootstrap validity and consistency of the resulting test, and show that the proposed procedure attains a double-robustness property with respect to conditional generator estimation errors. Simulations and real data applications suggest that the proposed method performs well for multivariate responses and high-dimensional covariates.
翻译:我们研究了利用生成模型检验两个条件分布是否相等的问题。所提出的方法从每个样本中学习一个条件生成器,并利用它生成在另一个样本中观测到的协变量值对应的响应,从而允许生成的响应与观测到的响应直接比较。通过交叉生成对齐协变量,该方法避免了条件密度比估计和高维协变量的局部平滑。该构造的总体版本产生一个条件差异量,在合适的重叠条件下表征两个条件分布相等性,而样本版本则导出一个检验统计量,定义为多重自助法校准的再生核希尔伯特空间索引经验过程的上确界。我们基于交替最大化和核技巧,开发了一种计算该统计量及其自助法模拟量的高效算法。在理论上,我们推导了原假设和备择假设下检验统计量的极限分布,证明了自助法的有效性和所得检验的一致性,并表明所提出的程序在条件生成器估计误差方面具有双重稳健性。模拟和真实数据应用表明,所提出的方法在多变量响应和高维协变量场景下表现良好。