Conditional independence (CI) is central to causal inference, feature selection, and graphical modeling, yet it is untestable in many settings without additional assumptions. Existing CI tests often rely on restrictive structural conditions, limiting their validity. Kernel methods using partial covariance operators offer a more principled approach but suffer from limited adaptivity and scalability. In this work, we explore whether representation learning can help address these limitations. Specifically, we focus on representations derived from the singular value decomposition of partial covariance operators and use them to construct a simple test statistic. We also introduce a bi-level contrastive algorithm to learn these representations. Our theory links representation learning error to test performance and establishes asymptotic validity and power guarantees. Experiments on real and synthetic data suggest that this approach offers a principled and statistically grounded path toward scalable CI testing, bridging kernel-based theory with modern representation learning.
翻译:条件独立性(CI)是因果推断、特征选择和图模型的核心概念,但在许多场景中若无额外假设则无法直接检验。现有CI检验方法通常依赖严格的结构性条件,限制了其有效性。基于部分协方差算子的核方法虽提供了更严谨的途径,但面临适应性与可扩展性不足的挑战。本研究探索表示学习能否突破这些限制:具体而言,我们聚焦于通过部分协方差算子的奇异值分解所获得的表示,并基于此构建简洁的检验统计量。同时,我们引入双层级对比学习算法来学习这些表示。理论分析揭示了表示学习误差与检验性能之间的关联,并建立了渐近有效性与检验功效的保证。在真实与合成数据上的实验表明,该方法为可扩展的CI检验提供了兼具理论严谨性与统计合理性的路径,有效桥接了基于核的理论与当代表示学习。