The task of moment localization is to localize a temporal moment in an untrimmed video for a given natural language query. Since untrimmed video contains highly redundant contents, the quality of the query is crucial for accurately localizing moments, i.e., the query should provide precise information about the target moment so that the localization model can understand what to look for in the videos. However, the natural language queries in current datasets may not be easy to understand for existing models. For example, the Ego4D dataset uses question sentences as the query to describe relatively complex moments. While being natural and straightforward for humans, understanding such question sentences are challenging for mainstream moment localization models like 2D-TAN. Inspired by the recent success of large language models, especially their ability of understanding and generating complex natural language contents, in this extended abstract, we make early attempts at reformulating the moment queries into a set of instructions using large language models and making them more friendly to the localization models.
翻译:时刻定位任务旨在根据给定的自然语言查询,从无修剪视频中定位出特定的时间片段。由于无修剪视频包含大量冗余内容,查询质量对于准确定位时刻至关重要——即查询应提供关于目标时刻的精确信息,以便定位模型理解在视频中寻找什么。然而,当前数据集中的自然语言查询可能难以被现有模型理解。例如,Ego4D数据集使用疑问句作为查询来描述相对复杂的时刻。虽然对人类而言这些疑问句自然且直观,但诸如2D-TAN等主流定位模型却难以理解此类问句。受大型语言模型近期成功(尤其是其理解和生成复杂自然语言内容的能力)的启发,本扩展摘要初步尝试利用大型语言模型将时刻查询重构为一组指令,使其对定位模型更加友好。