Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings. While recent work identifies these most influential sets, there is no formal way to tell when maximum influence is excessive rather than expected under natural random sampling variation. We address this gap by developing a principled framework for most influential sets. Focusing on linear least-squares, we derive a convenient exact influence formula and identify the extreme value distributions of maximal influence - the heavy-tailed Fréchet for constant-size sets and heavy-tailed data, and the well-behaved Gumbel for growing sets or light tails. This allows us to conduct rigorous hypothesis tests for excessive influence. We demonstrate through applications across economics, biology, and machine learning benchmarks, resolving contested findings and replacing ad-hoc heuristics with rigorous inference.


翻译:小的有影响力数据子集可能对模型结论产生巨大影响,少数数据点即可颠覆关键发现。尽管近期研究识别了这些最具影响力的集合,但尚未有正式方法能判断最大影响力何时为过度,而非自然随机采样变异性下的预期结果。我们通过开发一个关于最具影响力集合的原理性框架来填补这一空白。聚焦于线性最小二乘法,我们推导出一个便捷的精确影响力公式,并识别出最大影响力的极值分布——对于固定大小的集合和重尾数据为厚尾的弗雷歇分布,而对于增长的集合或轻尾数据则为良性的冈贝尔分布。这使得我们能够对过度影响力进行严格的假设检验。通过在经济学、生物学和机器学习基准测试中的应用,我们解决了具有争议的发现,并用严格的推断取代了临时性的启发式方法。

0
下载
关闭预览

相关内容

清华大学《《SuperBench大模型综合能力评测报告》发布
专知会员服务
47+阅读 · 2024年4月20日
事件抽取的再评价:过去、现在和未来的挑战
专知会员服务
25+阅读 · 2023年11月28日
【ICML2023】面向影响力最大化的深度图表示学习与优化
专知会员服务
29+阅读 · 2023年5月6日
CVPR 二十年,影响力最大的 10 篇论文!
专知会员服务
31+阅读 · 2022年2月1日
ISWC2020最佳论文《可解释假信息检测的链接可信度评价》
最新SCI期刊影响因子出炉!
CVer
27+阅读 · 2020年7月1日
Attention!注意力机制模型最新综述
中国人工智能学会
18+阅读 · 2019年4月8日
异常检测的阈值,你怎么选?给你整理好了...
机器学习算法与Python学习
10+阅读 · 2018年9月19日
【资源】史上最全数据集汇总
七月在线实验室
18+阅读 · 2018年4月24日
福利 | 最全面超大规模数据集下载链接汇总
AI研习社
26+阅读 · 2017年9月7日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Testing Preferential Sampling
Arxiv
0+阅读 · 6月12日
Arxiv
0+阅读 · 6月10日
Arxiv
0+阅读 · 5月7日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
6+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关VIP内容
清华大学《《SuperBench大模型综合能力评测报告》发布
专知会员服务
47+阅读 · 2024年4月20日
事件抽取的再评价:过去、现在和未来的挑战
专知会员服务
25+阅读 · 2023年11月28日
【ICML2023】面向影响力最大化的深度图表示学习与优化
专知会员服务
29+阅读 · 2023年5月6日
CVPR 二十年,影响力最大的 10 篇论文!
专知会员服务
31+阅读 · 2022年2月1日
ISWC2020最佳论文《可解释假信息检测的链接可信度评价》
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员