Ensemble methods combine the predictions of several base models. We study whether or not including more models in an ensemble always improve its average performance. Such a question depends on the kind of ensemble considered, as well as the predictive metric chosen. We focus on situations where all members of the ensemble are a priori expected to perform as well, which is the case of several popular methods like random forests or deep ensembles. In this setting, we essentially show that ensembles are getting better all the time if, and only if, the considered loss function is convex. More precisely, in that case, the average loss of the ensemble is a decreasing function of the number of models. When the loss function is nonconvex, we show a series of results that can be summarised by the insight that ensembles of good models keep getting better, and ensembles of bad models keep getting worse. To this end, we prove a new result on the monotonicity of tail probabilities that may be of independent interest. We illustrate our results on a simple machine learning problem (diagnosing melanomas using neural nets).
翻译:集成方法将多个基模型的预测结果进行组合。我们研究在集成中纳入更多模型是否总能提升其平均性能。这一问题取决于所考虑的集成类型以及所选用的预测指标。我们聚焦于集成中所有成员先验预期表现均等的场景——这正是随机森林、深度集成等几种流行方法所呈现的情形。在此设定下,我们本质上证明:当且仅当所考虑的损失函数为凸函数时,集成的性能才会持续提升。更精确地说,在此情况下,集成的平均损失是模型数量的递减函数。当损失函数非凸时,我们通过一系列可归纳为以下洞见的结果表明:优秀模型的集成会持续改善,而劣质模型的集成则会持续恶化。为此,我们证明了一个关于尾部概率单调性的新结论,该结论可能具有独立研究价值。我们通过一个简单的机器学习问题(使用神经网络诊断黑色素瘤)对上述结果进行了验证。