While data-driven confounder selection requires careful consideration, it is frequently employed in observational studies to adjust for confounding factors. Widely recognized criteria for confounder selection include the minimal set approach, which involves selecting variables relevant to both treatment and outcome, and the union set approach, which involves selecting variables for either treatment or outcome. These approaches are often implemented using heuristics and off-the-shelf statistical methods, where the degree of uncertainty may not be clear. In this paper, we focus on the false discovery rate (FDR) to measure uncertainty in confounder selection. We define the FDR specific to confounder selection and propose methods based on the mirror statistic, a recently developed approach for FDR control that does not rely on p-values. The proposed methods are free from p-values and require only the assumption of some symmetry in the distribution of the mirror statistic. It can be easily combined with sparse estimation and other methods that involve difficulties in deriving p-values. The properties of the proposed method are investigated by exhaustive numerical experiments. Particularly in high-dimensional data scenarios, our method outperforms conventional methods.
翻译:尽管数据驱动的混杂因素选择需要谨慎考虑,但在观察性研究中常被用于调整混杂因子。公认的混杂因素选择标准包括最小集合法(选取与处理变量和结局变量均相关的变量)和并集法(选取与处理变量或结局变量相关的变量)。这些方法通常采用启发式策略和现成统计方法实现,但其不确定性程度可能不明确。本文聚焦于用错误发现率(FDR)衡量混杂因素选择中的不确定性。我们定义了针对混杂因素选择的错误发现率,并提出了基于镜像统计量的方法,这是一种不依赖p值的错误发现率控制新方法。所提方法无需p值,仅需镜像统计量分布存在某种对称性假设,可便捷地与稀疏估计及其他难以推导p值的方法结合。通过详尽的数值实验验证了该方法性质,尤其在处理高维数据时,本方法优于传统方法。