High-dimensional variable selection, with many more covariates than observations, is widely documented in standard regression models, but there are still few tools to address it in non-linear mixed-effects models where data are collected repeatedly on several individuals. In this work, variable selection is approached from a Bayesian perspective and a selection procedure is proposed, combining the use of a spike-and-slab prior and the Stochastic Approximation version of the Expectation Maximisation (SAEM) algorithm. Similarly to Lasso regression, the set of relevant covariates is selected by exploring a grid of values for the penalisation parameter. The SAEM approach is much faster than a classical MCMC (Markov chain Monte Carlo) algorithm and our method shows very good selection performances on simulated data. Its flexibility is demonstrated by implementing it for a variety of nonlinear mixed effects models. The usefulness of the proposed method is illustrated on a problem of genetic markers identification, relevant for genomic-assisted selection in plant breeding.
翻译:高维变量选择(即协变量数量远多于观测数量)在标准回归模型中已有广泛研究,但在非线性混合效应模型(数据来自多个个体的重复测量)中仍缺乏相关工具。本研究从贝叶斯视角探讨变量选择问题,提出一种结合spike-and-slab先验分布与随机近似期望最大化(SAEM)算法的选择流程。与Lasso回归类似,通过遍历惩罚参数网格值来筛选相关协变量集。该SAEM方法比经典MCMC(马尔可夫链蒙特卡洛)算法快得多,且在模拟数据中展现出优异的选择性能。通过将其应用于多种非线性混合效应模型,验证了该方法的灵活性。最后以植物育种中基因组辅助选择的遗传标记识别问题为例,展示了所提方法的实用性。