In this work, we analyze the conditions under which information about the context of an input $X$ can improve the predictions of deep learning models in new domains. Following work in marginal transfer learning in Domain Generalization (DG), we formalize the notion of context as a permutation-invariant representation of a set of data points that originate from the same domain as the input itself. We offer a theoretical analysis of the conditions under which this approach can, in principle, yield benefits, and formulate two necessary criteria that can be easily verified in practice. Additionally, we contribute insights into the kind of distribution shifts for which the marginal transfer learning approach promises robustness. Empirical analysis shows that our criteria are effective in discerning both favorable and unfavorable scenarios. Finally, we demonstrate that we can reliably detect scenarios where a model is tasked with unwarranted extrapolation in out-of-distribution (OOD) domains, identifying potential failure cases. Consequently, we showcase a method to select between the most predictive and the most robust model, circumventing the well-known trade-off between predictive performance and robustness.
翻译:本文分析了在何种条件下,关于输入$X$的上下文信息能够提升深度学习模型在新领域的预测性能。沿袭领域泛化(DG)中边际迁移学习的研究工作,我们将上下文概念形式化为与输入同域数据点集合的置换不变表征。我们对该方法在理论上可能产生收益的条件进行了理论分析,并提出了两个可在实践中轻松验证的必要准则。此外,我们揭示了边际迁移学习能够保证鲁棒性的分布偏移类型。实证分析表明,我们的准则能够有效区分有利与不利场景。最后,我们证明可以可靠地检测到模型在分布外(OOD)域中被迫进行不合理外推的场景,从而识别潜在的失败案例。基于此,我们展示了一种在最具预测能力与最鲁棒模型之间进行选择的方法,规避了预测性能与鲁棒性之间众所周知的权衡困境。