Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings. While recent work identifies these most influential sets, there is no formal way to tell when maximum influence is excessive rather than expected under natural random sampling variation. We address this gap by developing a principled framework for most influential sets. Focusing on linear least-squares, we derive a convenient exact influence formula and identify the extreme value distributions of maximal influence - the heavy-tailed Fréchet for constant-size sets and heavy-tailed data, and the well-behaved Gumbel for growing sets or light tails. This allows us to conduct rigorous hypothesis tests for excessive influence. We demonstrate through applications across economics, biology, and machine learning benchmarks, resolving contested findings and replacing ad-hoc heuristics with rigorous inference.
翻译:小的有影响力数据子集可能对模型结论产生巨大影响,少数数据点即可颠覆关键发现。尽管近期研究识别了这些最具影响力的集合,但尚未有正式方法能判断最大影响力何时为过度,而非自然随机采样变异性下的预期结果。我们通过开发一个关于最具影响力集合的原理性框架来填补这一空白。聚焦于线性最小二乘法,我们推导出一个便捷的精确影响力公式,并识别出最大影响力的极值分布——对于固定大小的集合和重尾数据为厚尾的弗雷歇分布,而对于增长的集合或轻尾数据则为良性的冈贝尔分布。这使得我们能够对过度影响力进行严格的假设检验。通过在经济学、生物学和机器学习基准测试中的应用,我们解决了具有争议的发现,并用严格的推断取代了临时性的启发式方法。