In small-batch scientific deployments, labeled target outcomes may be too scarce for reliable shift estimation even when unlabeled target inputs are available. We address the complementary setting where the practitioner has a pre-specified label-shift correction from domain knowledge and asks whether incoming labeled outcomes support it. We show that the per-observation likelihood ratio between a label-shift-corrected predictive and the source predictive is a conditional e-value, so its running product is a nonnegative martingale and Ville's inequality yields an anytime-valid confirmation rule. The log martingale equals the cumulative negative log-predictive density (NLPD) gap between the source and the corrected predictive, converting routine model monitoring into a formal sequential test. Rejection means the incoming data support the posited correction relative to the source predictive, but it is not a precise estimate of the degree of shift. Closed forms are available for GP sources with Gaussian label-shift ratios. GP regression simulations validate Type I control, finite-sample power, miscalibration sensitivity, and the small-batch advantage of a reliable prior over label-based re-estimation.
翻译:在小批量科学应用中,即使有未标记的目标输入可用,标记的目标结果也可能过于稀缺而无法进行可靠的偏移估计。我们针对另一种场景:实践者根据领域知识预先指定了标签偏移修正,并询问传入的标记结果是否支持该修正。我们证明,标签偏移修正后的预测模型与源预测模型之间每个观测值的似然比是一个条件e值,因此其运行乘积是一个非负鞅,并且Ville不等式可产生一个随时有效的确认规则。对数鞅等于源预测模型与修正后预测模型之间的累积负对数预测密度(NLPD)差距,从而将常规模型监控转化为一个正式的序贯检验。拒绝零假设意味着传入数据支持相对于源预测模型所假设的修正,但这并非对偏移程度的精确估计。对于具有高斯标签偏移比例的高斯过程(GP)源模型,可得到闭合形式的解。GP回归模拟验证了第一类错误控制、有限样本功效、校准误差敏感性,以及基于可靠先验的标签重估计在小批量场景中的优势。