Model-X knockoffs is a flexible wrapper method for high-dimensional regression algorithms, which provides guaranteed control of the false discovery rate (FDR). Due to the randomness inherent to the method, different runs of model-X knockoffs on the same dataset often result in different sets of selected variables, which is undesirable in practice. In this paper, we introduce a methodology for derandomizing model-X knockoffs with provable FDR control. The key insight of our proposed method lies in the discovery that the knockoffs procedure is in essence an e-BH procedure. We make use of this connection, and derandomize model-X knockoffs by aggregating the e-values resulting from multiple knockoff realizations. We prove that the derandomized procedure controls the FDR at the desired level, without any additional conditions (in contrast, previously proposed methods for derandomization are not able to guarantee FDR control). The proposed method is evaluated with numerical experiments, where we find that the derandomized procedure achieves comparable power and dramatically decreased selection variability when compared with model-X knockoffs.
翻译:模型X knockoffs是一种用于高维回归算法的灵活包装方法,能够保证对错误发现率(FDR)的控制。由于该方法固有的随机性,在同一数据集上多次运行模型X knockoffs通常会得到不同的变量选择结果,这在实际应用中并不理想。本文提出了一种具有可证明FDR控制能力的去随机化模型X knockoffs方法。该方法的核心发现是,knockoffs过程本质上是一种e-BH过程。我们利用这一关联,通过聚合多次knockoff实现得到的e值来对模型X knockoffs进行去随机化。我们证明了该去随机化过程能在无额外条件下(相反,先前提出的去随机化方法无法保证FDR控制)将FDR控制在期望水平。通过数值实验评估,我们发现与模型X knockoffs相比,所提出的去随机化方法在保持相当统计功效的同时,显著降低了选择变异性。