The performance of learning models often deteriorates when deployed in out-of-sample environments. To ensure reliable deployment, we propose a stability evaluation criterion based on distributional perturbations. Conceptually, our stability evaluation criterion is defined as the minimal perturbation required on our observed dataset to induce a prescribed deterioration in risk evaluation. In this paper, we utilize the optimal transport (OT) discrepancy with moment constraints on the \textit{(sample, density)} space to quantify this perturbation. Therefore, our stability evaluation criterion can address both \emph{data corruptions} and \emph{sub-population shifts} -- the two most common types of distribution shifts in real-world scenarios. To further realize practical benefits, we present a series of tractable convex formulations and computational methods tailored to different classes of loss functions. The key technical tool to achieve this is the strong duality theorem provided in this paper. Empirically, we validate the practical utility of our stability evaluation criterion across a host of real-world applications. These empirical studies showcase the criterion's ability not only to compare the stability of different learning models and features but also to provide valuable guidelines and strategies to further improve models.
翻译:学习模型在非样本环境中部署时,其性能往往会下降。为确保可靠部署,我们提出了一种基于分布扰动的稳定性评估准则。从概念上讲,该稳定性评估准则定义为:在观测数据集上诱发预定程度的风险评估性能退化所需的最小扰动。本文利用带矩约束的最优传输(OT)差异度量在(样本, 密度)空间上量化该扰动,从而使得稳定性评估准则能够同时应对实际场景中最常见的两类分布偏移——数据损坏与子群体漂移。为提升实用价值,我们针对不同损失函数类别,提出了一系列易于处理的凸优化公式及计算方法。实现该目标的关键技术工具是本文提供的一对强对偶定理。通过涵盖多个实际应用场景的实证研究,我们验证了所提稳定性评估准则的实用性。这些实验表明,该准则不仅能比较不同学习模型与特征的稳定性,还能为模型优化提供有价值的指导策略。