Distributional Random Forest (DRF) is a flexible forest-based method to estimate the full conditional distribution of a multivariate output of interest given input variables. In this article, we introduce a variable importance algorithm for DRFs, based on the well-established drop and relearn principle and MMD distance. While traditional importance measures only detect variables with an influence on the output mean, our algorithm detects variables impacting the output distribution more generally. We show that the introduced importance measure is consistent, exhibits high empirical performance on both real and simulated data, and outperforms competitors. In particular, our algorithm is highly efficient to select variables through recursive feature elimination, and can therefore provide small sets of variables to build accurate estimates of conditional output distributions.
翻译:分布随机森林(DRF)是一种基于森林的灵活方法,用于在给定输入变量条件下估计多元输出的完整条件分布。本文基于成熟的"舍弃与重学"原则和MMD距离,提出了一种针对DRF的变量重要性算法。传统重要性度量仅能检测影响输出均值的变量,而本文提出的算法能更广泛地检测影响输出分布的变量。我们证明了所提重要性度量具有一致性,在真实数据和模拟数据上均表现出较高的实证性能,并优于其他竞争方法。特别地,该算法通过递归特征消除进行变量筛选时效率极高,因此能够提供小规模变量集来构建输出条件分布的精确估计。