We demonstrate how supervised learning can be decomposed into a two-stage procedure, where (1) all model parameters are selected in an unsupervised manner, and (2) the outputs y are added to the model, without changing the parameter values. This is achieved by a new model selection criterion that - in contrast to cross-validation - can be used also without access to y. For linear ridge regression, we bound the asymptotic out-of-sample risk of our method in terms of the optimal asymptotic risk. We also demonstrate that versions of linear and kernel ridge regression, smoothing splines, k-nearest neighbors, random forests, and neural networks, trained without access to y, perform similarly to their standard y-based counterparts. Hence, our results suggest that the difference between supervised and unsupervised learning is less fundamental than it may appear.
翻译:我们证明了监督学习可以分解为一个两阶段过程:(1)所有模型参数均以无监督方式选择,(2)在保持参数值不变的情况下将输出y加入到模型中。这一发现基于一种新型模型选择准则——与交叉验证不同,该准则无需访问y即可使用。对于线性岭回归,我们以最优渐进风险为基准,界定了该方法渐进样本外风险的上界。我们还证明,在没有访问y的情况下训练的线性/核岭回归、光滑样条、k近邻、随机森林和神经网络等模型,其表现与基于y的标准对应物相近。因此,我们的结果表明监督学习与无监督学习之间的差异可能并不像表面上看起来那么本质。