While model selection is a well-studied topic in parametric and nonparametric regression or density estimation, selection of possibly high-dimensional nuisance parameters in semiparametric problems is far less developed. In this paper, we propose a selective machine learning framework for making inferences about a finite-dimensional functional defined on a semiparametric model, when the latter admits a doubly robust estimating function and several candidate machine learning algorithms are available for estimating the nuisance parameters. We introduce a new selection criterion aimed at bias reduction in estimating the functional of interest based on a novel definition of pseudo-risk inspired by the double robustness property. Intuitively, the proposed criterion selects a pair of learners with the smallest pseudo-risk, so that the estimated functional is least sensitive to perturbations of a nuisance parameter. We establish an oracle property for a multi-fold cross-validation version of the new selection criterion which states that our empirical criterion performs nearly as well as an oracle with a priori knowledge of the pseudo-risk for each pair of candidate learners. Finally, we apply the approach to model selection of a semiparametric estimator of average treatment effect given an ensemble of candidate machine learners to account for confounding in an observational study which we illustrate in simulations and a data application.
翻译:虽然模型选择在参数和非参数回归或密度估计中是一个研究充分的课题,但在半参数问题中选择可能高维的干扰参数则远未成熟。本文提出一个选择性机器学习框架,用于对定义在半参数模型上的有限维泛函进行推断,其中该模型允许存在双重鲁棒估计函数,并且有多种候选机器学习算法可用于估计干扰参数。我们引入一个新的选择准则,旨在通过基于双重鲁棒性启发的伪风险新定义来减少感兴趣泛函估计的偏差。直观上,该准则选择伪风险最小的学习者对,使得估计泛函对干扰参数的扰动最不敏感。我们为该新选择准则的多重交叉验证版本建立了预言家性质,表明我们的经验准则的表现几乎与预先知道每对候选学习者的伪风险的预言家一样好。最后,我们将该方法应用于平均处理效应的半参数估计量模型选择,采用候选机器学习集成模型来调整观察性研究中的混杂因素,并通过模拟和实际数据应用进行验证。