The problems of Lasso regression and optimal design of experiments share a critical property: their optimal solutions are typically \emph{sparse}, i.e., only a small fraction of the optimal variables are non-zero. Therefore, the identification of the support of an optimal solution reduces the dimensionality of the problem and can yield a substantial simplification of the calculations. It has recently been shown that linear regression with a \emph{squared} $\ell_1$-norm sparsity-inducing penalty is equivalent to an optimal experimental design problem. In this work, we use this equivalence to derive safe screening rules that can be used to discard inessential samples. Compared to previously existing rules, the new tests are much faster to compute, especially for problems involving a parameter space of high dimension, and can be used dynamically within any iterative solver, with negligible computational overhead. Moreover, we show how an existing homotopy algorithm to compute the regularization path of the lasso method can be reparametrized with respect to the squared $\ell_1$-penalty. This allows the computation of a Bayes $c$-optimal design in a finite number of steps and can be several orders of magnitude faster than standard first-order algorithms. The efficiency of the new screening rules and of the homotopy algorithm are demonstrated on different examples based on real data.
翻译:Lasso回归与实验最优设计问题共享一个关键性质:它们的最优解通常是\textit{稀疏的},即仅有少量最优变量非零。因此,识别最优解的支持集可降低问题维度,并显著简化计算。近期研究表明,具有\textit{二次}$\ell_1$范数稀疏诱导惩罚的线性回归等价于一个实验最优设计问题。本文利用这种等价性推导出安全筛选规则,用于剔除无关样本。与现有规则相比,新测试的计算速度显著提升,尤其适用于高维参数空间问题,且可在任何迭代求解器中动态使用,计算开销可忽略。此外,我们展示了如何将用于计算lasso方法正则化路径的现有同伦算法,通过二次$\ell_1$惩罚进行重新参数化。这使得贝叶斯$c$-最优设计能在有限步内完成计算,其速度可比标准一阶算法快数个数量级。基于真实数据的多个示例验证了新筛选规则与同伦算法的效率。