Learning from weak, proxy, or relative supervision is common when ground-truth labels are unavailable, but robustness under distribution shift remains poorly understood because the supervision mechanism itself may change across environments. We formalize this phenomenon as supervision drift, defined as changes in $P(y \mid x, c)$ across contexts, and study it in CRISPR-Cas13d transcriptomic perturbation experiments where guide efficacy is inferred indirectly from RNA-seq responses. Using publicly available data spanning two human cell lines and multiple post-induction timepoints, we construct a controlled non-IID benchmark with explicit domain (cell line) and temporal shifts, while reusing a fixed weak-label construction across all contexts to avoid changing targets. Across linear and tree-based models, weak supervision supports meaningful learning in-domain (ridge $R^2 = 0.356$, Spearman $ρ= 0.442$) and partial cross-cell-line transfer ($ρ\approx 0.40$). In contrast, temporal transfer collapses across all model classes considered, yielding negative $R^2$ and weak or near-zero $ρ$ (ridge $R^2 = -0.145$, $ρ= 0.008$; XGBoost $R^2 = -0.155$, $ρ= 0.056$; random forest $R^2 = -0.322$, $ρ= 0.139$). Additional robustness analyses using externally recomputed weak labels, shift-score quantification, and simple mitigation baselines preserve the same qualitative pattern. Feature-label association and feature-importance analyses remain relatively stable across cell lines but change sharply over time, indicating that failures arise from supervision drift rather than model capacity or simple covariate shift. These results show that strong in-domain performance under weak supervision can be misleading and motivate feature stability as a lightweight diagnostic for non-transferability before deployment.


翻译:当真实标签不可用时,通过弱监督、代理监督或相对监督进行学习很常见,但在分布偏移下的鲁棒性仍未被充分理解,因为监督机制本身可能在不同环境中发生变化。我们将此现象形式化为监督漂移,即上下文间$P(y \mid x, c)$的变化,并在CRISPR-Cas13d转录组扰动实验中加以研究,其中引导效率通过RNA-seq响应间接推断。利用跨越两种人类细胞系和多个诱导后时间点的公开可用数据,我们构建了一个受控制的非独立同分布基准,包含显式域(细胞系)和时间偏移,同时在所有上下文中复用固定的弱标签构建以避免目标变化。在线性和基于树的模型中,弱监督支持有意义的域内学习(岭回归$R^2 = 0.356$,斯皮尔曼$ρ= 0.442$)和部分跨细胞系迁移($ρ\approx 0.40$)。相反,所有考虑模型类的时间迁移均失效,产生负$R^2$和弱或接近零的$ρ$(岭回归$R^2 = -0.145$,$ρ= 0.008$;XGBoost $R^2 = -0.155$,$ρ= 0.056$;随机森林$R^2 = -0.322$,$ρ= 0.139$)。使用外部重计算的弱标签、偏移分数量化及简单缓解基线的附加鲁棒性分析保持了相同的定性模式。特征-标签关联和特征重要性分析在细胞系间保持相对稳定,但随时间急剧变化,表明失败源于监督漂移而非模型容量或简单协变量偏移。这些结果表明,弱监督下的强域内性能可能具有误导性,并激发特征稳定性作为部署前非可迁移性的轻量级诊断指标。

0
下载
关闭预览

相关内容

【NeurIPS2023】用几何协调对抗表示学习视差
专知会员服务
27+阅读 · 2023年10月28日
【CMU博士论文】分布偏移下的不确定性量化,226页pdf
专知会员服务
31+阅读 · 2023年9月30日
【NeurIPS 2021】实例依赖的偏标记学习
专知会员服务
11+阅读 · 2021年11月28日
专知会员服务
28+阅读 · 2021年8月24日
专知会员服务
38+阅读 · 2021年3月29日
最新《弱监督预训练语言模型微调》报告,52页ppt
专知会员服务
38+阅读 · 2020年12月26日
【CVPR2019】弱监督图像分类建模
深度学习大讲堂
38+阅读 · 2019年7月25日
「PPT」深度学习中的不确定性估计
专知
27+阅读 · 2019年7月20日
你的算法可靠吗? 神经网络不确定性度量
专知
40+阅读 · 2019年4月27日
机器学习中如何处理不平衡数据?
机器之心
13+阅读 · 2019年2月17日
论文浅尝 | 基于深度强化学习的远程监督数据集的降噪
开放知识图谱
29+阅读 · 2019年1月17日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
VIP会员
最新内容
面向2027年及未来的海军情报改革
专知会员服务
2+阅读 · 8月5日
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
4+阅读 · 8月5日
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 8月5日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员