Temporal classification errors are often treated as representation failures, but they can also arise from how available evidence is converted into decisions. This paper proposes a representation--calibration decomposition for temporal classification. We keep a trained native classifier frozen and separate two inference-time interventions: a conservative residual multi-scale branch that adds auxiliary logits to the native prediction, and a post-hoc branch-aware calibrator that recombines native and residual evidence at decision time. This design distinguishes missing temporal evidence from underused decision-level evidence without retraining the backbone. Across FI-2010, PTB-XL, UCI-HAR, MHEALTH, and HARTH, we find that gains are strongly regime-dependent. Residual multi-scale evidence is most useful in noisy or representation-limited settings, especially short-horizon FI-2010 and weaker recurrent backbones, while branch-aware calibration helps when native and auxiliary logits contain complementary evidence not fully exploited by the raw decision rule. Near-saturated settings show limited gains from either intervention. These results suggest that temporal classification should be understood not only as representation learning, but also as the problem of trusting, combining, and calibrating evidence from multiple views.
翻译:时序分类错误常被视为表征学习的失败,但同样可能源于将可用证据转化为决策的方式。本文提出一种针对时序分类的表征-校准分解方法。我们保持预训练的原生分类器不变,分离两种推理时干预:一是保守的残差多尺度分支,用于向原生预测添加辅助逻辑值;二是事后分支感知校准器,在决策阶段重组原生证据与残差证据。这种设计无需重新训练主干网络即可区分缺失的时序证据与未被充分利用的决策级证据。在FI-2010、PTB-XL、UCI-HAR、MHEALTH和HARTH数据集上的实验表明,性能提升具有强烈的任务依赖性:残差多尺度证据在噪声或表征受限场景(尤其是短时域FI-2010数据集及较弱循环神经网络主干)中最具效用;而分支感知校准在原生逻辑值与辅助逻辑值包含互补性证据且未被原始决策规则充分利用时发挥作用。接近饱和的场景下两种干预均收益有限。这些结果表明,时序分类不仅应被视为表征学习问题,更应理解为多视角证据的信任、组合与校准问题。