Existing deep video models are limited by specific tasks, fixed input-output spaces, and poor generalization capabilities, making it difficult to deploy them in real-world scenarios. In this paper, we present our vision for multimodal and versatile video understanding and propose a prototype system, \system. Our system is built upon a tracklet-centric paradigm, which treats tracklets as the basic video unit and employs various Video Foundation Models (ViFMs) to annotate their properties e.g., appearance, motion, \etc. All the detected tracklets are stored in a database and interact with the user through a database manager. We have conducted extensive case studies on different types of in-the-wild videos, which demonstrates the effectiveness of our method in answering various video-related problems. Our project is available at https://www.wangjunke.info/ChatVideo/
翻译:现有深度视频模型受限于特定任务、固定的输入输出空间以及较差的泛化能力,难以部署于真实场景。本文提出了我们对多模态通用视频理解的愿景,并构建了一个原型系统\system。该系统基于以轨迹片段为中心的设计范式,将轨迹片段作为视频的基本单元,并利用多种视频基础模型(ViFMs)标注其属性(如外观、运动等)。所有检测到的轨迹片段存储于数据库中,通过数据库管理器与用户交互。我们在多种类型的野生视频上开展了广泛的案例研究,结果证明了本方法在回答各类视频相关问题上的有效性。项目代码见 https://www.wangjunke.info/ChatVideo/