Linear mixed models (LMMs), which incorporate fixed and random effects, are key tools for analyzing heterogeneous data, such as in personalized medicine. Nowadays, this type of data is increasingly wide, sometimes containing thousands of candidate predictors, necessitating sparsity for prediction and interpretation. However, existing sparse learning methods for LMMs do not scale well beyond tens or hundreds of predictors, leaving a large gap compared with sparse methods for linear models, which ignore random effects. This paper closes the gap with a new $\ell_0$ regularized method for LMM subset selection that can run on datasets containing thousands of predictors in seconds to minutes. On the computational front, we develop a coordinate descent algorithm as our main workhorse and provide a guarantee of its convergence. We also develop a local search algorithm to help traverse the nonconvex optimization surface. Both algorithms readily extend to subset selection in generalized LMMs via a penalized quasi-likelihood approximation. On the statistical front, we provide a finite-sample bound on the Kullback-Leibler divergence of the new method. We then demonstrate its excellent performance in experiments involving synthetic and real datasets.
翻译:线性混合模型(LMMs)通过融合固定效应和随机效应,成为分析异质性数据(如个性化医疗)的关键工具。当前此类数据日益呈现高维特征,有时包含数千个候选预测变量,因此需通过稀疏化方法提升预测精度与可解释性。然而,现有针对LMMs的稀疏学习方法在预测变量超过数十或数百个后难以扩展,与忽略随机效应的线性模型稀疏方法存在显著差距。本文提出一种新型$\ell_0$正则化子集选择方法填补这一空白,该方法可在数秒至数分钟内处理包含数千个预测变量的数据集。在计算方面,我们以坐标下降算法为核心工作引擎并提供其收敛性保证,同时开发局部搜索算法辅助遍历非凸优化曲面。上述两种算法通过惩罚拟似然近似可自然扩展至广义线性混合模型的子集选择。在统计方面,我们推导了新方法Kullback-Leibler散度的有限样本界。实验证明,该方法在合成数据集与真实数据集上均展现卓越性能。