Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested. We examine one concrete instance: the use of routing statistics in Mixture-of-Experts (MoE) pruning, where utilization rates, activation norms, and routing weight distributions are treated as predictors of which experts can be removed without functional cost. A token-level interventional audit across three high-redundancy MoE architectures (OLMoE-1B-7B-0924, Qwen1.5-MoE-A2.7B, DeepSeek-V2-Lite) finds no observational metric predicts causal expert importance in any model: across all 60 metric-layer combinations effect sizes stay below Cohen's $d = 0.23$, and no metric is reliably positive under our corrected, dual-test criterion. A per-token routing weight control, run with identical $n$, rules out insufficient power, recovering a signal whose CI excludes zero at OLMoE's final MoE layer ($d = +0.231$, 95\% CI $[+0.09, +0.37]$, $p = 0.0013$). Existing pruning methods succeed in this regime not by identifying dispensable experts but because early-layer redundancy renders most selection criteria interchangeable. Our results provide an explicit counterexample to the common inferential step from population-level observational summaries to token-level interventional claims about expert importance, and illustrate how interventional audits can calibrate the evidential standards for interpretability claims.
翻译:可解释性方法通常使用基于观测模型行为的总体统计量,来推断特定干预对具体计算的影响;用珀尔术语表示,即把第一层关联性证据视为支持第二层干预结论,但这种推断的有效性鲜少被检验。我们考察了一个具体案例:混合专家模型(MoE)剪枝中的路由统计量使用——将利用率、激活范数和路由权重分布视为判断哪些专家可被移除而不影响功能的预测指标。对三个高冗余MoE架构(OLMoE-1B-7B-0924、Qwen1.5-MoE-A2.7B、DeepSeek-V2-Lite)进行的Token级干预审计发现,没有任何观测指标能预测任一模型中的因果专家重要性:在所有60个指标-层级组合中,效应量均低于Cohen's $d = 0.23$,且在我们经校正的双重检验标准下,无指标能稳定呈现正值。通过同样本量的逐Token路由权重控制,排除了统计效力不足的可能性,恢复出OLMoE最终MoE层上排除零值的信号($d = +0.231$,95% CI [+0.09, +0.37],$p = 0.0013$)。现有剪枝方法在此场景下成功并非因识别出可舍弃的专家,而是因为早期层冗余性使得多数选择指标可互换。我们的结果为“从总体观测统计量推断Token级专家重要性干预结论”这一常见推演步骤提供了明确反例,并展示了干预审计如何校准可解释性主张的证据标准。