Comparing Variable Selection and Model Averaging Methods for Logistic Regression

Model uncertainty is a central challenge in statistical models for binary outcomes such as logistic regression, arising when it is unclear which predictors should be included in the model. Many methods have been proposed to address this issue for logistic regression, but their relative performance under realistic conditions remains poorly understood. We therefore conducted a preregistered, simulation-based comparison of 28 established methods for variable selection and inference under model uncertainty, using 11 empirical datasets spanning a range of sample sizes and number of predictors, in cases both with and without separation. We found that Bayesian model averaging (BMA) methods based on g-priors, particularly g = max(n, p^2), show the strongest overall performance when separation is absent. When separation occurs, penalized likelihood approaches, especially the LASSO, provide the most stable results, while BMA with the local empirical Bayes (EB-local) prior is competitive in both situations. These findings offer practical guidance for applied researchers on how to effectively address model uncertainty in logistic regression in modern empirical and machine learning research.

翻译：模型不确定性是逻辑回归等二值结果统计模型中的一个核心挑战，其源于不确定应将哪些预测变量纳入模型。尽管已有多种方法被提出用于解决逻辑回归中的这一问题，但它们在现实条件下的相对表现仍不明确。为此，我们基于预注册设计，通过模拟研究比较了28种在模型不确定性下处理变量选择与推断的成熟方法，使用11个涵盖不同样本量和预测变量数量的实证数据集（包括存在和不存在分离情况）。研究发现：当不存在分离时，基于g先验的贝叶斯模型平均方法（特别是g = max(n, p^2)）整体表现最优；当存在分离时，惩罚似然方法（尤以LASSO）提供最稳定的结果，而采用局部经验贝叶斯先验的BMA方法在两种情境下均具竞争力。这些发现为应用研究者在现代实证研究与机器学习中有效应对逻辑回归模型不确定性提供了实用指导。

相关内容

逻辑回归

关注 318

逻辑回归（也称“对数几率回归”）（英语：Logistic regression 或logit regression），即逻辑模型（英语：Logit model，也译作“评定模型”、“分类评定模型”）是离散选择法模型之一，属于多重变量分析范畴，是社会学、生物统计学、临床、数量心理学、计量经济学、市场营销等统计实证分析的常用方法。在统计学中，logistic模型(或logit模型)用于对存在的某个类或事件的概率建模，例如通过/失败、赢/输、活着/死了或健康/生病。这可以扩展到建模若干类事件，如确定一个图像是否包含猫、狗、狮子等。图像中检测到的每个物体的概率都在0到1之间，其和为1。

《鲁棒优化中保形预测生成不确定性集的性能评价》最新95页

专知会员服务

10+阅读 · 3月20日

《不确定条件下优化问题的高效精确与近似算法》MIT最新130页

专知会员服务

31+阅读 · 2025年11月19日