Transfer learning is usually studied as a consequence of distribution shift. This paper identifies an orthogonal failure mode in which the data distribution is fixed and the loss changes. This setting is called \emph{loss shift}. A loss determines which information in \(X\) is Bayes-relevant, and two losses may therefore require different representations even under the same joint law \(P(X,Y)\). The idea is formalized using Bayes quotients, which allow losses to be ordered by refinement. In the Bayes-quotient formulation, strict refinement gives an immediate qualitative obstruction. A source-minimal representation for a coarser loss is insufficient for a strictly finer target loss. For finite-output log loss, this obstruction becomes an exact quantitative identity. The excess risk is the conditional information about \(Y\) discarded by the representation. Experiments in controlled, learned, synthetic-image, and real-image settings show the predicted effect, i.e., classification-equivalent representations can have different optimal log-loss performance under a fixed data distribution.
翻译:迁移学习通常被视为分布偏移的结果。本文识别了一种正交的失效模式,其中数据分布固定而损失函数发生变化。这种设置被称为“损失迁移”。损失函数决定了 \(X\) 中哪些信息与贝叶斯相关,因此即使在相同的联合分布 \(P(X,Y)\) 下,不同的损失函数也可能需要不同的表示。该思想通过贝叶斯商数形式化,允许损失函数按精细化程度排序。在贝叶斯商数框架中,严格精细化会立即产生定性障碍:针对较粗糙损失函数的源最小表示不足以保证严格更精细化的目标损失函数。对于有限输出的对数损失,这种障碍转化为精确的定量恒等式,超额风险即为表示丢弃的关于 \(Y\) 的条件信息。在受控、学习、合成图像和真实图像实验设置中均观测到预测效应:即使数据分布固定,分类等价的表示可能具有不同的最优对数损失性能。