In this paper, error estimates of classification Random Forests are quantitatively assessed. Based on the initial theoretical framework built by Bates et al. (2023), the true error rate and expected error rate are theoretically and empirically investigated in the context of a variety of error estimation methods common to Random Forests. We show that in the classification case, Random Forests' estimates of prediction error is closer on average to the true error rate instead of the average prediction error. This is opposite the findings of Bates et al. (2023) which were given for logistic regression. We further show that this result holds across different error estimation strategies such as cross-validation, bagging, and data splitting.
翻译:本文对分类随机森林的误差估计进行了定量评估。基于Bates等人(2023年)构建的初始理论框架,我们从理论和实证角度研究了随机森林常见误差估计方法下的真实错误率和期望错误率。研究表明,在分类问题中,随机森林对预测误差的估计平均更接近真实错误率而非平均预测误差。这一结果与Bates等人(2023年)针对逻辑回归提出的结论相反。我们进一步证明,该结论在交叉验证、自助聚合和数据拆分等不同误差估计策略下均成立。