Predictive modelling is often reduced to finding the best model that optimizes a selected performance measure. But what if the second-best model describes the data equally well but in a completely different way? What about the third? Is it possible that the most effective models learn completely different relationships in the data? Inspired by Anscombe's quartet, this paper introduces Rashomon's quartet, a synthetic dataset for which four models from different classes have practically identical predictive performance. However, their visualization reveals drastically distinct ways of understanding the correlation structure in data. The introduced simple illustrative example aims to further facilitate visualization as a mandatory tool to compare predictive models beyond their performance. We need to develop insightful techniques for the explanatory analysis of model sets.
翻译:预测建模通常被简化为选择最能优化特定性能指标的模型。然而,若次优模型对数据的描述同样准确但完全以不同方式呢?第三优模型呢?是否可能存在这样的情况:最有效的模型从数据中学习到了截然不同的关系?受安斯库姆四重奏启发,本文引入了拉什莫尔四重奏——一个合成数据集,其中四个来自不同类别的模型具有几乎相同的预测性能。然而,它们的可视化揭示了在理解数据相关结构时存在截然不同的方式。这个简单示例旨在进一步促进可视化作为超越性能比较预测模型所必需的通用工具。我们亟需开发深入技术,用于对模型集合进行解释性分析。