Long video understanding is a significant and ongoing challenge in the intersection of multimedia and artificial intelligence. Employing large language models (LLMs) for comprehending video becomes an emerging and promising method. However, this approach incurs high computational costs due to the extensive array of video tokens, experiences reduced visual clarity as a consequence of token aggregation, and confronts challenges arising from irrelevant visual tokens while answering video-related questions. To alleviate these issues, we present an Interactive Visual Adapter (IVA) within LLMs, designed to enhance interaction with fine-grained visual elements. Specifically, we first transform long videos into temporal video tokens via leveraging a visual encoder alongside a pretrained causal transformer, then feed them into LLMs with the video instructions. Subsequently, we integrated IVA, which contains a lightweight temporal frame selector and a spatial feature interactor, within the internal blocks of LLMs to capture instruction-aware and fine-grained visual signals. Consequently, the proposed video-LLM facilitates a comprehensive understanding of long video content through appropriate long video modeling and precise visual interactions. We conducted extensive experiments on nine video understanding benchmarks and experimental results show that our interactive visual adapter significantly improves the performance of video LLMs on long video QA tasks. Ablation studies further verify the effectiveness of IVA in long and short video understandings.
翻译:长视频理解是多媒体与人工智能交叉领域中的一个重要且持续存在的挑战。利用大语言模型理解视频成为一种新兴且有前景的方法。然而,该方法因大量视频令牌而产生高昂的计算成本,因令牌聚合而导致视觉清晰度降低,并在回答视频相关问题时会面临无关视觉令牌带来的挑战。为解决这些问题,我们在LLMs中提出了一种交互式视觉适配器,旨在增强对细粒度视觉元素的交互能力。具体而言,我们首先利用视觉编码器与预训练的因果变换器将长视频转换为时序视频令牌,随后将其与视频指令一同输入LLMs。接着,我们在LLMs内部模块中集成了IVA,该适配器包含轻量级时序帧选择器和空间特征交互器,以捕获指令感知且细粒度的视觉信号。因此,所提出的视频-LLM通过适当的长视频建模和精准视觉交互,促进了对长视频内容的全面理解。我们在九个视频理解基准上进行了广泛实验,结果表明,我们的交互式视觉适配器显著提升了视频LLMs在长视频问答任务上的性能。消融研究进一步验证了IVA在长视频与短视频理解中的有效性。