Classical results establish that ensembles of small models benefit when predictive diversity is encouraged, through bagging, boosting, and similar. Here we demonstrate that this intuition does not carry over to ensembles of deep neural networks used for classification, and in fact the opposite can be true. Unlike regression models or small (unconfident) classifiers, predictions from large (confident) neural networks concentrate in vertices of the probability simplex. Thus, decorrelating these points necessarily moves the ensemble prediction away from vertices, harming confidence and moving points across decision boundaries. Through large scale experiments, we demonstrate that diversity-encouraging regularizers hurt the performance of high-capacity deep ensembles used for classification. Even more surprisingly, discouraging predictive diversity can be beneficial. Together this work strongly suggests that the best strategy for deep ensembles is utilizing more accurate, but likely less diverse, component models.
翻译:经典研究结果表明,通过装袋、提升等方法鼓励预测多样性有助于小型模型的集成。但本文证明,这一直觉并不适用于用于分类的深度神经网络集成,实际上可能恰恰相反。与回归模型或小型(不自信)分类器不同,大型(自信)神经网络的预测结果集中于概率单纯形的顶点附近。因此,对这些点去相关化必然导致集成预测偏离顶点,既损害置信度又使预测点跨越决策边界。通过大规模实验,我们证明鼓励多样性的正则化项会损害用于分类的高容量深度集成的性能。更令人惊讶的是,抑制预测多样性反而可能有益。综上,本研究强烈表明,深度集成的最佳策略是采用更准确、但可能多样性较低的组件模型。