Video diffusion models exhibit emergent reasoning capabilities like solving mazes and puzzles, yet little is understood about how they reason during generation. We take a first step towards understanding this and study the internal planning dynamics of video models using 2D maze solving as a controlled testbed. Our investigations reveal two findings. Our first finding is early plan commitment: video diffusion models commit to a high-level motion plan within the first few denoising steps, after which further denoising alters visual details but not the underlying trajectory. Our second finding is that path length, not obstacle density, is the dominant predictor of maze difficulty, with a sharp failure threshold at 12 steps. This means video models can only reason over long mazes by chaining together multiple sequential generations. To demonstrate the practical benefits of our findings, we introduce Chaining with Early Planning, or ChEaP, which only spends compute on seeds with promising early plans and chains them together to tackle complex mazes. This improves accuracy from 7% to 67% on long-horizon mazes and by 2.5x overall on hard tasks in Frozen Lake and VR-Bench across Wan2.2-14B and HunyuanVideo-1.5. Our analysis reveals that current video models possess deeper reasoning capabilities than previously recognized, which can be elicited more reliably with better inference-time scaling.
翻译:视频扩散模型展现出类似解决迷宫和谜题的浮现推理能力,但关于它们在生成过程中如何推理的理解却很少。我们首次尝试理解这一过程,并以二维迷宫求解作为受控测试平台,研究视频模型的内部规划动态。我们的调查揭示了两项发现。第一项发现是早期计划承诺:视频扩散模型在前几步去噪过程中就承诺了一个高层运动计划,后续去噪步骤只会改变视觉细节,而不会改变底层轨迹。第二项发现是路径长度(而非障碍物密度)是迷宫难度的主要预测指标,并且在12步时出现急剧的失败阈值。这意味着视频模型只能通过将多个连续生成片段链接起来,在长迷宫中推理。为了展示我们发现的实用价值,我们引入了早规划链接(ChEaP,Chaining with Early Planning),它仅将计算资源分配给那些具有前景早期计划的种子,并将它们链接起来以应对复杂迷宫。在长时域迷宫中,该方法将准确率从7%提升至67%,在Frozen Lake和VR-Bench的困难任务中,在Wan2.2-14B和HunyuanVideo-1.5两个模型上整体提升2.5倍。我们的分析表明,当前视频模型具备比先前认知更深的推理能力,且通过更好的推理时扩展可以更可靠地激发这种能力。