This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via MultimodAl Retrieval (MAGMaR). Addressing the critical challenges of cross-lingual long-video comprehension, strict persona adherence, and zero-hallucination temporal grounding, we propose a fully training-free, two-stage cascaded Video RAG pipeline. Our architecture strategically decouples semantic retrieval from cognitive logical reasoning through a modality-aware division of labor. In the first stage, a high-recall semantic pre-fetching module employs dense retrieval using only high-fidelity visual summaries and global text descriptions, explicitly isolating noisy modalities (e.g., OCR and ASR) to maintain a pristine vector space. In the second stage, an Adaptive, Iterative, and Reasoning-based (A.I.R.) filtering agent, powered by a commercial Large Language Model (LLM), performs fine-grained cognitive reranking. The agent re-incorporates full multimodal contexts to enforce strict logical alignment with user personas, effectively pruning semantically similar but logically irrelevant candidates. Finally, a Prompt Sculpting mechanism constrains the generator to synthesize the distilled subset into strictly formatted JSON responses with exact chunk-level citations. Evaluated on the RAG track, our resource-aware approach shows exceptional precision in both information retrieval and persona-conditioned generation.


翻译:本文介绍了我们为第二届通过多模态检索增强多模态生成研讨会(MAGMaR)提交的系统方案。针对跨语言长视频理解、严格角色一致性遵循以及零幻觉时间定位的关键挑战,我们提出了一种完全无需训练、采用两级级联架构的视频RAG流水线。我们的架构通过模态感知的职责分工,从策略上将语义检索与认知逻辑推理解耦。在第一阶段,一个高召回率的语义预取模块仅使用高保真视觉摘要和全局文本描述进行稠密检索,明确隔离噪声模态(如OCR和ASR)以保持纯净的向量空间。在第二阶段,一个基于自适应、迭代与推理(A.I.R.)的过滤智能体(由商业大语言模型驱动)执行细粒度认知重排序。该智能体重整完整的多模态上下文,强制实现与用户角色的严格逻辑对齐,有效剔除语义相似但逻辑无关的候选内容。最后,提示塑形机制约束生成器将蒸馏得到的子集综合为严格格式化的JSON响应,并附以精确的片段级引用。在RAG赛道上的评估表明,我们这种资源感知的方法在信息检索和角色条件化生成两个任务上均展现出卓越的精度。

0
下载
关闭预览

相关内容

【SIGIR2025教程】动态与参数化检索增强生成
专知会员服务
17+阅读 · 2025年7月14日
多模态检索增强生成的综合综述
专知会员服务
44+阅读 · 2025年2月17日
图增强生成(GraphRAG)
专知会员服务
35+阅读 · 2025年1月4日
【CVPR2024】OmniViD: 一个用于通用视频理解的生成框架
专知会员服务
25+阅读 · 2024年3月27日
视频文本预训练简述
专知会员服务
22+阅读 · 2022年7月24日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
VIP会员
最新内容
《履带式无人地面战车技术发展现状》
专知会员服务
4+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
4+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
12+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
7+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
10+阅读 · 7月31日
相关VIP内容
【SIGIR2025教程】动态与参数化检索增强生成
专知会员服务
17+阅读 · 2025年7月14日
多模态检索增强生成的综合综述
专知会员服务
44+阅读 · 2025年2月17日
图增强生成(GraphRAG)
专知会员服务
35+阅读 · 2025年1月4日
【CVPR2024】OmniViD: 一个用于通用视频理解的生成框架
专知会员服务
25+阅读 · 2024年3月27日
视频文本预训练简述
专知会员服务
22+阅读 · 2022年7月24日
相关资讯
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员