The performance of machine learning models can be impacted by changes in data over time. A promising approach to address this challenge is invariant learning, with a particular focus on a method known as invariant risk minimization (IRM). This technique aims to identify a stable data representation that remains effective with out-of-distribution (OOD) data. While numerous studies have developed IRM-based methods adaptive to data augmentation scenarios, there has been limited attention on directly assessing how well these representations preserve their invariant performance under varying conditions. In our paper, we propose a novel method to evaluate invariant performance, specifically tailored for IRM-based methods. We establish a bridge between the conditional expectation of an invariant predictor across different environments through the likelihood ratio. Our proposed criterion offers a robust basis for evaluating invariant performance. We validate our approach with theoretical support and demonstrate its effectiveness through extensive numerical studies.These experiments illustrate how our method can assess the invariant performance of various representation techniques.
翻译:机器学习模型的性能可能因数据随时间变化而受到影响。应对这一挑战的有效方法是不变学习,其中重点关注一种称为不变风险最小化(IRM)的方法。该技术旨在识别一种稳定的数据表示,使其在分布外(OOD)数据上仍能保持有效。尽管已有大量研究开发了适用于数据增强场景的IRM方法,但针对这些表示在变化条件下如何维持其不变性能的直接评估仍鲜有关注。本文提出了一种专为IRM方法设计的不变性能评估新方法。我们通过似然比建立了不同环境下不变预测器条件期望之间的桥梁。所提出的准则为评估不变性能提供了稳健的基础。我们通过理论支撑验证了该方法的有效性,并通过大量数值实验展示了其评估多种表示技术不变性能的能力。