A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |W_{out} W_{in}|$, with early experiments suggesting the phenomenon generalized across architectures and benefited from interventional data at $p < 10^{-5}$. We package the protocol used to test that claim -- standardized synthetic generators (VAR/Lorenz/CauseMe-style), three intervention semantics ($do(X=c)$, soft-noise, random-forcing), edge-provenance cards on three real datasets, and size-matched control arms -- as a reusable falsification benchmark, and walk the claim through it in five stages. The method-level claim does not survive: (i) a plain linear bottleneck does as well or better; (ii) tuned Lasso beats the bottleneck on synthetic CauseMe-style benchmarks, and on Lorenz-96 (the only real benchmark with unambiguous ground truth) classical PCMCI and Granger lead a tight cluster in which the bottleneck trails; (iii) the headline intervention advantage is roughly 60% a sample-size confound, and the residual disappears under standard $do(X=c)$ interventions, surviving only under a non-standard random-forcing scheme; (iv) even that residual reproduces, with a larger effect, in classical bivariate Granger -- the effect is method-agnostic. What survives is a narrow characterization result; the benchmark is the lasting artifact, and each stage above is one of its control arms.
翻译:仅针对下一步预测训练的Mamba状态空间模型,似乎能通过简单读取操作$S=|W_{out}W_{in}|$恢复格兰杰因果结构。早期实验表明,该现象可跨架构泛化,并在$p<10^{-5}$的干预数据中受益。我们将验证该主张的标准化协议体系——包括标准合成数据生成器(VAR/Lorenz/CauseMe风格)、三种干预语义($do(X=c)$、软噪声、随机强迫)、三个真实数据集上的边缘溯源卡片以及规模匹配的对照组——打包为可复用的证伪基准,分五个阶段对该主张进行验证。方法层面的主张未能成立:(i)普通线性瓶颈模型表现相当或更优;(ii)调优的Lasso在合成CauseMe风格基准和Lorenz-96系统(唯一具有明确真实值的真实基准)上均优于瓶颈模型,而经典PCMCI和格兰杰检验在此处形成紧密聚类,瓶颈模型落后于该聚类;(iii)标题所示的干预优势中约60%源于样本量混杂,剩余部分在标准$do(X=c)$干预下消失,仅存在非标准随机强迫方案中;(iv)该剩余效应在经典双变量格兰杰检验中复现且效应更大——表明该效应与具体方法无关。最终保留的仅为狭隘的表征结论;本研究所建立的基准体系才是持久性成果,而上述每个阶段均构成其对照组之一。