Action reasoning in egocentric video requires capturing fine-grained transitions of hand-object interactions, a task where general-purpose Vision-Language Models (VLMs) often struggle when operating directly on raw pixels. We propose to decouple visual perception from symbolic reasoning by converting videos into Temporal Action Graphs. In a multi-stage prompting pipeline, we first generate dense natural language narratives over short temporal windows as a semantic bottleneck, then formalize them into structured, open-vocabulary graph representations. On the EGTEA and Epic-Kitchens-100 datasets, the symbolic representation unlocks efficient in-context learning: few-shot graph demonstrations yield substantial accuracy gains over zero-shot frame and graph-based inference alike. Even in the zero-shot setting, graph-based reasoning remains competitive with pixel-based inference despite potential pretraining contamination favoring the latter. Across 11 open-weight VLMs from 6 model families ranging from 2B to 235B parameters, our findings indicate that current VLMs are more effective as symbolic reasoners than as direct visual observers. By projecting video into the language domain, we provide a scalable, fine-tuning-free alternative to end-to-end approaches that better leverages these models' latent reasoning strengths. The code will be made public.


翻译:自我中心视频中的动作推理需要捕捉手-物交互的细粒度过渡,而通用视觉语言模型(VLMs)在直接处理原始像素时通常难以胜任此类任务。我们提出通过将视频转换为时序动作图谱,将视觉感知与符号推理解耦。在多阶段提示流水线中,我们首先在短时窗口内生成密集的自然语言叙事作为语义瓶颈,随后将其形式化为结构化的开放词汇图谱表示。在EGTEA和Epic-Kitchens-100数据集上,符号表示实现了高效的上下文学习:小样本图谱演示相较零样本帧和图谱推理均取得了显著精度提升。即使在零样本设置中,基于图谱的推理仍能与基于像素的推理保持竞争力,尽管后者可能存在预训练污染优势。跨越6个模型家族(参数规模从2B到235B)的11个开放权重VLM的实验表明,当前VLM作为符号推理器的有效性优于直接视觉观察器。通过将视频投影到语言域,我们为端到端方法提供了一种可扩展、免微调的替代方案,能够更充分地利用这些模型的潜在推理能力。代码将开源。

0
下载
关闭预览

相关内容

综述 | 自我中心视频中的视觉语言模型
专知会员服务
6+阅读 · 8月21日
在无标注条件下适配视觉—语言模型:全面综述
专知会员服务
13+阅读 · 2025年8月9日
大规模视觉-语言模型的基准、评估、应用与挑战
专知会员服务
18+阅读 · 2025年2月10日
基于大语言模型的时序知识图谱推理模型蒸馏方法
专知会员服务
38+阅读 · 2025年1月10日
《遥感时序视觉语言模型》全面综述
专知会员服务
30+阅读 · 2024年12月4日
【NeurIPS2023】大型语言模型是视觉推理协调器
专知会员服务
30+阅读 · 2023年10月24日
一文看懂如何将深度学习应用于视频动作识别
ETP:精确时序动作定位
极市平台
13+阅读 · 2018年5月25日
实战 | 基于深度学习模型VGG的图像识别(附代码)
七月在线实验室
13+阅读 · 2018年3月30日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
刚刚!Jev中文教程项目发布了
专知会员服务
0+阅读 · 10月4日
《人工智能赋能的适应性多功能电磁战》
专知会员服务
12+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
8+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
11+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
17+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
11+阅读 · 9月23日
相关VIP内容
综述 | 自我中心视频中的视觉语言模型
专知会员服务
6+阅读 · 8月21日
在无标注条件下适配视觉—语言模型:全面综述
专知会员服务
13+阅读 · 2025年8月9日
大规模视觉-语言模型的基准、评估、应用与挑战
专知会员服务
18+阅读 · 2025年2月10日
基于大语言模型的时序知识图谱推理模型蒸馏方法
专知会员服务
38+阅读 · 2025年1月10日
《遥感时序视觉语言模型》全面综述
专知会员服务
30+阅读 · 2024年12月4日
【NeurIPS2023】大型语言模型是视觉推理协调器
专知会员服务
30+阅读 · 2023年10月24日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员