Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe this gap with a deceptively simple but revealing challenge: judging the arrow of time (AoT)-whether a short clip is played forward or backward. We introduce AoT-PsyPhyBENCH, a psychophysically validated benchmark that tests whether VLMs can infer temporal direction in natural videos using the same stimuli and behavioral baselines established for humans. Our comprehensive evaluation of open-weight and proprietary, reasoning and non-reasoning VLMs reveals that most models perform near chance, and even the best model lags far behind human accuracy on physically irreversible processes (e.g., free fall, diffusion/explosion) and causal manual actions (division/addition) that humans recognize almost instantly. These results highlight a fundamental gap in current multimodal systems: while they capture rich visual-semantic correlations, they lack the inductive biases required for temporal continuity and causal understanding. We release the code and data for AoT-PsyPhyBENCH to encourage further progress in the physical and temporal reasoning capabilities of VLMs.
翻译:现代视觉-语言模型(VLM)在多模态任务中表现卓越,但对视频中时间信息的理解仍然薄弱且缺乏充分评估。我们通过一个看似简单却极具揭示性的挑战——判断时间之箭(AoT),即判断短视频片段是正向播放还是反向播放——来探究这一差距。为此,我们构建了AoT-PsyPhyBENCH基准,该基准采用经心理物理学验证的评估方法,基于与人类相同的刺激素材和行为基线标准,测试VLM是否能在自然视频中推断时间方向。我们对开源模型与商业模型、推理模型与非推理模型进行了全面评估,结果显示:绝大多数模型的表现接近随机水平,即使性能最优的模型,在处理人类几乎能瞬间识别的物理不可逆过程(如自由落体、扩散/爆炸)及因果性操作行为(如拆分/组合)时,其准确率仍远低于人类。这些结果揭示了当前多模态系统的根本性缺陷:虽然它们能捕捉丰富的视觉-语义关联,但缺乏时间连续性和因果理解所需的归纳偏置。我们已开源AoT-PsyPhyBENCH的代码与数据,以期推动VLM在物理推理与时间推理能力方面的进一步发展。