This work addresses the problem of automated covariate selection under limited prior knowledge. Given an exposure-outcome pair {X,Y} and a variable set Z of unknown causal structure, the Local Discovery by Partitioning (LDP) algorithm partitions Z into subsets defined by their relation to {X,Y}. We enumerate eight exhaustive and mutually exclusive partitions of any arbitrary Z and leverage this taxonomy to differentiate confounders from other variable types. LDP is motivated by valid adjustment set identification, but avoids the pretreatment assumption commonly made by automated covariate selection methods. We provide theoretical guarantees that LDP returns a valid adjustment set for any Z that meets sufficient graphical conditions. Under stronger conditions, we prove that partition labels are asymptotically correct. Total independence tests is worst-case quadratic in |Z|, with sub-quadratic runtimes observed empirically. We numerically validate our theoretical guarantees on synthetic and semi-synthetic graphs. Adjustment sets from LDP yield less biased and more precise average treatment effect estimates than baselines, with LDP outperforming on confounder recall, test count, and runtime for valid adjustment set discovery.
翻译:本文研究在有限先验知识下自动协变量选择的问题。给定暴露-结果对{X,Y}以及一个因果结构未知的变量集Z,分区局部发现(LDP)算法将Z划分为由它们与{X,Y}关系定义的子集。我们枚举了任意Z的八种完备且互斥的分区,并利用这一分类体系区分混杂因子与其他变量类型。LDP以有效调整集识别为动机,但避免了自动协变量选择方法通常采用的预处理假设。我们提供了理论保证:对于任何满足充分图条件的Z,LDP都能返回一个有效的调整集。在更强条件下,我们证明了分区标签的渐近正确性。总独立性检验次数在最坏情况下与|Z|呈二次关系,而经验观察到的运行时间低于二次。我们在合成图和半合成图上对理论保证进行了数值验证。与基线方法相比,LDP得出的调整集产生的平均处理效应估计偏差更小、精度更高,且在混杂因子召回率、检验次数以及有效调整集发现运行时间方面均表现更优。