Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect genuine spatial intelligence. We show that existing spatial VLMs lack basic camera motion understanding, a key component of spatial cognition. We propose the Spatial Narrative Score (SNS), an evaluation framework that requires VLMs to generate explicit spatial narratives capturing both scene semantics and camera motion, followed by reasoning with a frozen proxy LLM. Under SNS, state-of-the-art spatial VLMs exhibit significant performance degradation despite high direct question answering accuracy. To address this gap, we introduce CaMo, a camera motion grounded VLM that achieves consistent performance across SNS evaluation and direct spatial question answering accuracy. Our results highlight the importance of explicit spatial narrative externalization for evaluating VLMs with transferable 3D spatial understanding. Our code, data, and model is available at https://github.com/hsiangwei0903/CaMo
翻译:视觉语言模型(VLM)在空间问答基准测试中取得了强劲的性能,但尚不明确这些提升是否反映了真正的空间智能。我们发现现有空间VLM缺乏基本的相机运动理解能力,而这是空间认知的关键组成部分。我们提出空间叙事分数(SNS)评估框架,该框架要求VLM生成同时包含场景语义和相机运动的显式空间叙事,随后通过冻结的代理大语言模型进行推理。在SNS框架下,最先进的空间VLM尽管在直接问答任务中表现高准确率,其性能却显著下降。为解决这一差距,我们提出CaMo——一种基于相机运动的VLM,在SNS评估与直接空间问答准确率上均实现一致性能。我们的结果强调了显式空间叙事外化对于评估具有可迁移三维空间理解能力的VLM的重要性。我们的代码、数据和模型已开源在https://github.com/hsiangwei0903/CaMo。