Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful causal reasoning. We introduce a dual-probe methodology that isolates these properties. The Text-Only Probe measures linguistic quality. The Chain-Text Probe requires models to first generate explicit causal chains. The Abstraction Gap (AG) metric quantifies the normalized performance difference. Evaluating eight VLMs on CAGE (Causal Abstraction Gap Evaluation), a benchmark of 49,500 questions across 5,500 images spanning Pearl's causal hierarchy, we find seven models exhibit AG exceeding 0.50 with text scores of 6--8 but chain scores below 2.5. Fine-tuning on 45,000 chain-annotated examples fails to close the gap. However, one model achieves near-zero AG. The capability exists within current VLM architectures and depends on pretraining and architectural choices. CAGE provides a diagnostic tool for assessing faithful causal reasoning in VLMs.
翻译:视觉-语言模型(VLM)能够生成流畅的因果解释,但当前的评估方法无法区分语言合理性与真正的因果推理能力。我们提出了一种双探针方法论来隔离这些特性。文本探针用于评估语言质量,而链式文本探针则要求模型首先生成显式的因果链。抽象鸿沟(AG)指标量化了归一化的性能差异。在CAGE(因果抽象鸿沟评估)基准上——该基准包含5,500张图像中的49,500个问题,覆盖Pearl因果层次——对八个VLM的评估显示,七个模型的AG值超过0.50,其文本得分在6-8之间,但链式得分低于2.5。对45,000个带链式注释的样本进行微调未能缩小这一鸿沟。然而,有一个模型实现了接近零的AG值。这一能力存在于当前的VLM架构中,并依赖于预训练和架构选择。CAGE为评估VLM中的忠实因果推理提供了诊断工具。