Variable selection methods are required in practical statistical modeling, to identify and include only the most relevant predictors, and then improving model interpretability. Such variable selection methods are typically employed in regression models, for instance in this article for the Poisson Log Normal model (PLN, Chiquet et al., 2021). This model aim to explain multivariate count data using dependent variables, and its utility was demonstrating in scientific fields such as ecology and agronomy. In the case of the PLN model, most recent papers focus on sparse networks inference through combination of the likelihood with a L1 -penalty on the precision matrix. In this paper, we propose to rely on a recent penalization method (SIC, O'Neill and Burke, 2023), which consists in smoothly approximating the L0-penalty, and that avoids the calibration of a tuning parameter with a cross-validation procedure. Moreover, this work focuses on the coefficient matrix of the PLN model and establishes an inference procedure ensuring effective variable selection performance, so that the resulting fitted model explaining multivariate count data using only relevant explanatory variables. Our proposal involves implementing a procedure that integrates the SIC penalization algorithm (epsilon-telescoping) and the PLN model fitting algorithm (a variational EM algorithm). To support our proposal, we provide theoretical results and insights about the penalization method, and we perform simulation studies to assess the method, which is also applied on real datasets.
翻译:在实际统计建模中,变量选择方法必不可少,用于识别并仅保留最相关的预测变量,从而提升模型可解释性。此类方法通常应用于回归模型,例如本文研究的泊松对数正态模型(PLN, Chiquet等, 2021)。该模型旨在利用解释变量解释多元计数数据,其效用已在生态学和农学等科学领域得到验证。针对PLN模型,近期文献主要关注通过似然函数与精度矩阵上的L1惩罚项相结合实现稀疏网络推断。本文提出采用一种新型惩罚方法(SIC, O'Neill and Burke, 2023),该方法通过平滑近似L0惩罚项,避免使用交叉验证程序校准调节参数。此外,本研究聚焦PLN模型的系数矩阵,建立了一套确保有效变量选择性能的推断流程,使最终拟合模型仅用相关解释变量即可解释多元计数数据。我们的方案通过集成SIC惩罚算法(ε-望远镜算法)与PLN模型拟合算法(变分EM算法)实现。为支撑该方案,我们提供了关于惩罚方法的理论结果与洞见,并通过模拟研究评估方法性能,同时将其应用于真实数据集。