Data valuation has become central in the era of data-centric AI. It drives efficient training pipelines and enables objective pricing in data markets by assigning a numeric value to each data point. Most existing data valuation methods estimate the effect of removing individual data points by evaluating changes in model validation performance under in-distribution (ID) settings, as opposed to out-of-distribution (OOD) scenarios where data follow different patterns. Since ID and OOD data behave differently, data valuation methods based on ID loss often fail to generalize to OOD settings, particularly when the validation set contains no OOD data. Furthermore, although OOD-aware methods exist, they involve heavy computational costs, which hinder practical deployment. To address these challenges, we introduce \emph{Eigen-Value} (EV), a plug-and-play data valuation framework for OOD robustness that uses only an ID data subset, including during validation. EV provides a new spectral approximation of domain discrepancy, which is the gap of loss between ID and OOD using ratios of eigenvalues of ID data's covariance matrix. EV then estimates the marginal contribution of each data point to this discrepancy via perturbation theory, alleviating the computational burden. Subsequently, EV plugs into ID loss-based methods by adding an EV term without any additional training loop. We demonstrate that EV achieves improved OOD robustness and stable value rankings across real-world datasets, while remaining computationally lightweight. These results indicate that EV is practical for large-scale settings with domain shift, offering an efficient path to OOD-robust data valuation.


翻译:摘要:数据估值已成为数据驱动人工智能时代的关键环节。它通过为每个数据点分配数值,推动高效训练流程并实现数据市场中的客观定价。现有的大多数数据估值方法通过评估在领域内(ID)设置下移除单个数据点对模型验证性能的影响来估算其效果,而非数据遵循不同分布的领域外(OOD)场景。由于领域内与领域外数据表现不同,基于领域内损失的数据估值方法通常难以推广至领域外设置,尤其是当验证集不包含领域外数据时。此外,尽管存在领域外感知方法,但其高昂的计算成本阻碍了实际部署。为解决这些挑战,我们提出了Eigen-Value(EV),一种即插即用的数据估值框架,仅需使用领域内数据子集(包括验证阶段)即可实现领域外鲁棒性。EV通过利用领域内数据协方差矩阵的特征值比值,提出了一种领域差异的光谱近似方法,该差异即领域内与领域外损失之间的差距。随后,EV基于扰动理论估算每个数据点对该差异的边际贡献,从而减轻计算负担。进一步,EV通过添加一个无需额外训练循环的EV项,可嵌入基于领域内损失的方法中。我们证明,EV在实际数据集上实现了更优的领域外鲁棒性和稳定的价值排序,同时保持计算轻量级。这些结果表明,EV适用于存在领域偏移的大规模场景,为领域外鲁棒的数据估值提供了一条高效路径。

0
下载
关闭预览

相关内容

国家标准《人工智能深度学习算法评估》(征求意见稿)
【MIT博士论文】实用机器学习的高效鲁棒算法,142页pdf
专知会员服务
60+阅读 · 2022年9月7日
美智库最新报告:小数据人工智能潜力不可估量,39页pdf
专知会员服务
77+阅读 · 2021年11月18日
专知会员服务
52+阅读 · 2021年7月14日
基于深度学习的数据融合方法研究综述
专知
37+阅读 · 2020年12月10日
你的算法可靠吗? 神经网络不确定性度量
专知
40+阅读 · 2019年4月27日
基于数据的分布式鲁棒优化算法及其应用【附PPT与视频资料】
人工智能前沿讲习班
27+阅读 · 2018年12月13日
【大数据】海量数据分析能力形成和大数据关键技术
产业智能官
17+阅读 · 2018年10月29日
【深度学习】深度学习的核心:掌握训练数据的方法
产业智能官
12+阅读 · 2018年1月14日
关于数据挖掘,有几本书推荐给你......
图灵教育
16+阅读 · 2017年10月11日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
0+阅读 · 24分钟前
俄乌战争中关于中程打击无人机部署的经验启示
专知会员服务
0+阅读 · 41分钟前
《基于强化学习的自动化红队测试》
专知会员服务
4+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
相关基金
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
25+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员