Long-video multimodal question answering requires structured reasoning over visual evidence and dialogue, but Large Vision-Language Models (LVLMs) are constrained by context-window and compute limits. We propose POVQA, which compresses each second into a temporally pooled image (1 fps pooled images) to maintain dense temporal coverage under a fixed token budget. We then train Qwen2.5-VL-7B with supervised fine-tuning (SFT) on rationale+answer targets, and optionally apply Direct Preference Optimization (DPO) for preference alignment. We introduce ReasonVQA as a pilot diagnostic dataset with 12 movies and 239 human-annotated QA+rationale triplets for controlled analysis of long-context multimodal reasoning under compression. On ReasonVQA, SFT improves the best pooled-only baseline from 0.212 to 0.550 F1, showing that pooled evidence plus rationale supervision provides the main performance gains in this setting. In zero-shot transfer, POVQA also reaches 64.7\% on TVQA after SFT+DPO. These results are preliminary: ReasonVQA is small, pooling can lose fine-grained temporal order, and DPO effects are not uniformly positive across settings. Code, dataset, and additional qualitative evaluations are available at \href{https://povqa.github.io}{https://povqa.github.io}.


翻译:长视频多模态问答需要对视觉证据和对话进行结构化推理,但大型视觉语言模型受限于上下文窗口和计算资源限制。我们提出POVQA方法,通过将每秒压缩为时间池化图像(1帧/秒的池化图像),在固定令牌预算下保持密集的时间覆盖。我们采用监督微调在推理+答案目标上训练Qwen2.5-VL-7B,并可选地应用直接偏好优化进行偏好对齐。我们引入ReasonVQA作为试点诊断数据集,包含12部电影及239个人工标注的问答+推理三元组,用于受控分析压缩条件下的长上下文多模态推理。在ReasonVQA上,监督微调将最佳纯池化基线从0.212提升至0.550 F1分数,表明池化证据与推理监督相结合在此场景下取得了主要性能增益。在零样本迁移中,POVQA在应用监督微调+直接偏好优化后于TVQA上达到64.7%准确率。这些结果尚属初步:ReasonVQA规模较小,池化可能丢失细粒度时间顺序,且直接偏好优化的效果在不同场景下并非一致正向。代码、数据集及更多定性评估结果请访问 \href{https://povqa.github.io}{https://povqa.github.io}。

0
下载
关闭预览

相关内容

思想来自于视觉机制,是对信息进行抽象的过程。
【CVPR2025】BIMBA:面向长范围视频问答的选择性扫描压缩
【CVPR2024】MoReVQA:探索视频问答的模块化推理模型
专知会员服务
18+阅读 · 2024年4月10日
【2022新书】视觉问答 (VQA):从理论到应用
专知会员服务
63+阅读 · 2022年5月24日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
44+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
相关资讯
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
44+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员