Machine learning regression methods allow estimation of functions without unrealistic parametric assumptions. Although they can perform exceptionally in prediction error, most lack theoretical convergence rates necessary for semi-parametric efficient estimation (e.g. TMLE, AIPW) of parameters like average treatment effects. The Highly Adaptive Lasso (HAL) is the only regression method proven to converge quickly enough for a meaningfully large class of functions, independent of the dimensionality of the predictors. Unfortunately, HAL is not computationally scalable. In this paper we build upon the theory of HAL to construct the Selectively Adaptive Lasso (SAL), a new algorithm which retains HAL's dimension-free, nonparametric convergence rate but which also scales computationally to large high-dimensional datasets. To accomplish this, we prove some general theoretical results pertaining to empirical loss minimization in nested Donsker classes. Our resulting algorithm is a form of gradient tree boosting with an adaptive learning rate, which makes it fast and trivial to implement with off-the-shelf software. Finally, we show that our algorithm retains the performance of standard gradient boosting on a diverse group of real-world datasets. SAL makes semi-parametric efficient estimators practically possible and theoretically justifiable in many big data settings.
翻译:机器学习回归方法可在无需不切实际的参数假设情况下估计函数。尽管这些方法在预测误差方面表现卓越,但大多数缺乏半参数高效估计(如TMLE、AIPW)诸如平均处理效应等参数所需的理论收敛速度。高适应LASSO(HAL)是唯一被证明能对有意义的大类函数实现快速收敛的回归方法,且该收敛速度独立于预测变量的维度。遗憾的是,HAL在计算上不可扩展。本文基于HAL理论构建了选择性自适应LASSO(SAL),这是一种既能保留HAL无维度非参数收敛速度,又能在大规模高维数据集上实现可扩展计算的新算法。为此,我们证明了嵌套Donsker类中经验损失最小化的一些通用理论结果。所提出的算法采用自适应学习率的梯度树提升形式,使其可通过现成软件快速实现且易于部署。最后,我们证明该算法在多样化的真实数据集上保持了标准梯度提升的性能。SAL使半参数高效估计器在实际可行性和理论合理性方面适用于众多大数据场景。