When variable selection methods are applied to bootstrapped and multiply imputed datasets, the set of selected variables typically varies across iterations. Aggregating results via the union rule can lead to overly dense models. We propose a sequential evidence aggregation procedure that models detection outcomes across perturbation iterations as Bernoulli trials and accumulates evidence for variable relevance through a likelihood-ratio process admitting an approximate Bayes-factor interpretation. The procedure provides both a variable inclusion criterion and a stopping rule that eliminates the need to fix the number of bootstrap-imputation iterations ex ante. A Monte Carlo study across 126 scenarios and an empirical illustration demonstrate the method's performance relative to existing aggregation approaches.
翻译:当变量选择方法应用于Bootstrap重抽样与多重插补数据集时,所选变量的集合通常在多次迭代间存在差异。通过取并集规则汇总结果可能导致模型过于稠密。本文提出一种序贯证据聚合方法,将扰动迭代中的检测结果建模为伯努利试验,并通过具有近似贝叶斯因子解释的似然比过程累积变量相关性的证据。该过程既提供了变量纳入准则,又给出了停止规则,从而无需事先固定Bootstrap-插补迭代次数。基于126种场景的蒙特卡洛研究与实证分析表明,该方法相较于现有聚合策略具有更优性能。