Predictive modelling is often reduced to finding the best model that optimizes a selected performance measure. But what if the second-best model describes the data in a completely different way? What about the third-best? Is it possible that the equally effective models describe different relationships in the data? Inspired by Anscombe's quartet, this paper introduces a Rashomon quartet, a four models built on synthetic dataset which have practically identical predictive performance. However, their visualization reveals distinct explanations of the relation between input variables and the target variable. The illustrative example aims to encourage the use of visualization to compare predictive models beyond their performance.
翻译:预测建模常常被简化为寻找优化选定性能指标的最佳模型。但若次优模型以截然不同的方式描述数据,又当如何?第三优的模型呢?同等有效的模型是否可能描述数据中不同的关系?受安斯库姆四重奏启发,本文引入了《罗生门四重奏》,即基于合成数据集构建的四个预测性能几乎完全相同的模型。然而,其可视化揭示了输入变量与目标变量之间关系的不同解释。这一示例旨在鼓励在比较预测模型时,使用可视化方法而非仅关注其性能。