While test-time scaling has revolutionized reasoning in large language models, generative video reasoning remains bottlenecked by a single-shot paradigm. We demonstrate that searching over denoising steps cannot rescue logically flawed rollouts because spatial trajectories commit early in the diffusion process. Root-level Best-of-N (BoN) sampling is similarly inefficient: reasoning errors cluster early in the temporal axis, and resampling blindly discards verified upstream progress. To unlock effective test-time scaling for video models, we introduce Temporal Backtracking Search (TBS), which shifts the search space to the temporal axis. TBS transforms video generation into an iterative generate-verify-restart loop via three core mechanisms: (1) variable-K conditioning to resume generation from arbitrary clean prefixes; (2) temporal process verification to localize failures and extract valid restart anchors; and (3) prefix-based search to reallocate compute toward extending correct trajectories rather than root resampling. Across algorithmic, navigation, and robotics domains, TBS Pareto-dominates matched-budget BoN. In a strict out-of-distribution setting where one-shot generation collapses (0.7% for BoN), TBS achieves 22.7%, with every solved episode stemming from a restarted branch. Ultimately, TBS reveals that the local reasoning competence of video models far exceeds what single-shot rollouts indicate, providing a scalable test-time framework to unlock it.


翻译:尽管测试时扩展已在大型语言模型中彻底革新了推理能力,但生成式视频推理仍受限于一次性推理范式。我们证明,在去噪步骤中进行搜索无法挽救逻辑有缺陷的展开,因为空间轨迹在扩散过程早期就已定型。根级最佳N选采样同样低效:推理误差沿时间轴早期聚集,而重采样会盲目丢弃已验证的上游进展。为解锁视频模型的有效测试时扩展,我们提出时间回溯搜索,将搜索空间迁移至时间轴。TBS通过三项核心机制将视频生成转化为迭代式生成-验证-重启动循环:(1)可变K条件化,能从任意干净前缀恢复生成;(2)时间过程验证,用于定位失败并提取有效重启动锚点;(3)基于前缀的搜索,将计算资源重新分配至扩展正确轨迹而非根重采样。在算法、导航和机器人领域,TBS帕累托支配了同等预算的BoN。在严格分布外场景中,一次性生成方法失效(BoN仅0.7%),而TBS达到22.7%,且每个已解决片段均源自重启分支。最终,TBS揭示了视频模型的局部推理能力远超一次性展开所表现的水平,为解锁该能力提供了可扩展的测试时框架。

0
下载
关闭预览

相关内容

VBVR:超大规模视频推理评测与数据集套件
专知会员服务
7+阅读 · 3月2日
基于扩散模型和流模型的推理时引导生成技术
专知会员服务
17+阅读 · 2025年4月30日
视频生成中的物理认知演进探究:一项综述
专知会员服务
17+阅读 · 2025年3月30日
生成技术在时空数据挖掘中的应用
专知会员服务
39+阅读 · 2024年6月5日
【AAAI2021】知识图谱增强的预训练模型的生成式常识推理
通过集成 XNNPACK 实现推理速度飞跃
TensorFlow
26+阅读 · 2020年7月30日
因果推理学习算法资源大列表
专知
27+阅读 · 2019年3月3日
基于视频的目标检测的发展【附PPT与视频资料】
人工智能前沿讲习班
19+阅读 · 2018年12月14日
关系推理:基于表示学习和语义要素
计算机研究与发展
19+阅读 · 2017年8月22日
回归预测&时间序列预测
GBASE数据工程部数据团队
44+阅读 · 2017年5月17日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
0+阅读 · 6月12日
VIP会员
最新内容
《履带式无人地面战车技术发展现状》
专知会员服务
4+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
4+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
12+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
10+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
6+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
9+阅读 · 7月31日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员