Distributional Random Forest (DRF) is a flexible forest-based method to estimate the full conditional distribution of a multivariate output of interest given input variables. In this article, we introduce a variable importance algorithm for DRFs, based on the well-established drop and relearn principle and MMD distance. While traditional importance measures only detect variables with an influence on the output mean, our algorithm detects variables impacting the output distribution more generally. We show that the introduced importance measure is consistent, exhibits high empirical performance on both real and simulated data, and outperforms competitors. In particular, our algorithm is highly efficient to select variables through recursive feature elimination, and can therefore provide small sets of variables to build accurate estimates of conditional output distributions.
翻译:分布随机森林(DRF)是一种灵活的基于森林的方法,用于在给定输入变量的条件下估计多元输出变量的完整条件分布。本文基于成熟的"剔除与重新学习"原则和MMD距离,提出了一种针对DRF的变量重要性算法。传统重要性度量仅能检测对输出均值有影响的变量,而我们的算法能更广泛地检测影响输出分布的变量。我们证明该重要性度量具有一致性,在真实数据和模拟数据上均展现出优异的实证性能,并优于竞争方法。特别地,该算法通过递归特征消除进行变量选择时效率极高,从而能够提供少量变量以构建条件输出分布的准确估计。