Temporal Sentence Grounding (TSG), which aims to localize moments from videos based on the given natural language queries, has attracted widespread attention. Existing works are mainly designed for short videos, failing to handle TSG in long videos, which poses two challenges: i) complicated contexts in long videos require temporal reasoning over longer moment sequences, and ii) multiple modalities including textual speech with rich information require special designs for content understanding in long videos. To tackle these challenges, in this work we propose a Grounding-Prompter method, which is capable of conducting TSG in long videos through prompting LLM with multimodal information. In detail, we first transform the TSG task and its multimodal inputs including speech and visual, into compressed task textualization. Furthermore, to enhance temporal reasoning under complicated contexts, a Boundary-Perceptive Prompting strategy is proposed, which contains three folds: i) we design a novel Multiscale Denoising Chain-of-Thought (CoT) to combine global and local semantics with noise filtering step by step, ii) we set up validity principles capable of constraining LLM to generate reasonable predictions following specific formats, and iii) we introduce one-shot In-Context-Learning (ICL) to boost reasoning through imitation, enhancing LLM in TSG task understanding. Experiments demonstrate the state-of-the-art performance of our Grounding-Prompter method, revealing the benefits of prompting LLM with multimodal information for TSG in long videos.
翻译:时间句子定位(Temporal Sentence Grounding,TSG)旨在根据自然语言查询从视频中定位相关时刻,已引起广泛关注。现有方法主要针对短视频设计,难以应对长视频中的TSG任务,这带来两大挑战:i) 长视频中的复杂语境需要对更长的时刻序列进行时间推理;ii) 包含丰富信息的文本语音等多模态数据需要特殊设计以理解长视频内容。为应对这些挑战,本文提出Grounding-Prompter方法,通过利用多模态信息提示大语言模型(LLM)来执行长视频中的TSG任务。具体而言,我们首先将TSG任务及其多模态输入(包括语音和视觉)转化为压缩的任务文本化形式。此外,为增强复杂语境下的时间推理,提出了一种边界感知提示策略,包含三个层面:i) 设计新型多尺度去噪思维链(CoT),逐步结合全局与局部语义并过滤噪声;ii) 建立有效性原则,约束LLM按特定格式生成合理预测;iii) 引入单样本上下文学习(ICL),通过模仿提升推理能力,增强LLM对TSG任务的理解。实验表明,我们的Grounding-Prompter方法取得了最先进性能,揭示了利用多模态信息提示LLM实现长视频TSG的优势。