Machine learning models in modern mass-market applications are often updated over time. One of the foremost challenges faced is that, despite increasing overall performance, these updates may flip specific model predictions in unpredictable ways. In practice, researchers quantify the number of unstable predictions between models pre and post update -- i.e., predictive churn. In this paper, we study this effect through the lens of predictive multiplicity -- i.e., the prevalence of conflicting predictions over the set of near-optimal models (the Rashomon set). We show how traditional measures of predictive multiplicity can be used to examine expected churn over this set of prospective models -- i.e., the set of models that may be used to replace a baseline model in deployment. We present theoretical results on the expected churn between models within the Rashomon set from different perspectives. And we characterize expected churn over model updates via the Rashomon set, pairing our analysis with empirical results on real-world datasets -- showing how our approach can be used to better anticipate, reduce, and avoid churn in consumer-facing applications. Further, we show that our approach is useful even for models enhanced with uncertainty awareness.
翻译:现代大众市场应用中的机器学习模型经常随时间更新。面临的主要挑战之一是,尽管更新提升了整体性能,但这些更新可能以不可预测的方式翻转特定模型的预测结果。实践中,研究人员量化更新前后模型间不稳定预测的数量——即预测性流失。本文通过预测性多重性视角研究这一现象——即近优模型集合(拉什蒙集)中冲突预测的普遍性。我们展示了如何利用预测性多重性的传统度量指标来考察预期模型更新集合的流失情况——即可能用于替代部署基线模型的模型集合。我们从不同角度提出了关于拉什蒙集内模型间预期流失的理论结果,并通过拉什蒙集表征了模型更新过程中的预期流失,结合真实数据集上的实证结果进行分析——证明我们的方法能更好地预测、减少并避免面向消费者应用中的流失。此外,我们证明该方法对增强不确定性感知的模型同样有效。