To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we firstly release the largest public Chinese high-quality video-language dataset named Youku-mPLUG, which is collected from Youku, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, and quality. Youku-mPLUG contains 10 million Chinese video-text pairs filtered from 400 million raw videos across a wide range of 45 diverse categories for large-scale pre-training. In addition, to facilitate a comprehensive evaluation of video-language models, we carefully build the largest human-annotated Chinese benchmarks covering three popular video-language tasks of cross-modal retrieval, video captioning, and video category classification. Youku-mPLUG can enable researchers to conduct more in-depth multimodal research and develop better applications in the future. Furthermore, we release popular video-language pre-training models, ALPRO and mPLUG-2, and our proposed modularized decoder-only model mPLUG-video pre-trained on Youku-mPLUG. Experiments show that models pre-trained on Youku-mPLUG gain up to 23.1% improvement in video category classification. Besides, mPLUG-video achieves a new state-of-the-art result on these benchmarks with 80.5% top-1 accuracy in video category classification and 68.9 CIDEr score in video captioning, respectively. Finally, we scale up mPLUG-video based on the frozen Bloomz with only 1.7% trainable parameters as Chinese multimodal LLM, and demonstrate impressive instruction and video understanding ability. The zero-shot instruction understanding experiment indicates that pretraining with Youku-mPLUG can enhance the ability to comprehend overall and detailed visual semantics, recognize scene text, and leverage open-domain knowledge.
翻译:为推动中文社区中视觉语言预训练(VLP)及多模态大语言模型(LLM)的发展,我们首次发布了最大规模的公开中文高质量视频语言数据集——Youku-mPLUG。该数据集取自国内知名视频分享网站优酷,严格遵循安全性、多样性与质量标准,从4亿原始视频中筛选出覆盖45个广泛类别的1000万个中文视频-文本对,可用于大规模预训练。此外,为促进视频语言模型的全面评估,我们精心构建了最大规模的人工标注中文基准测试集,涵盖跨模态检索、视频描述与视频分类三大主流视频语言任务。Youku-mPLUG将使研究者能够开展更深入的多模态研究,并开发出更优的未来应用。同时,我们发布了流行的视频语言预训练模型ALPRO与mPLUG-2,以及基于Youku-mPLUG预训练的自研模块化解码器模型mPLUG-video。实验结果表明,基于Youku-mPLUG预训练的模型在视频分类任务中性能提升高达23.1%。此外,mPLUG-video在上述基准测试中分别以80.5%的视频分类Top-1准确率与68.9的CIDEr分数(视频描述任务)刷新了当时最优结果。最终,我们基于冻结的Bloomz模型仅以1.7%可训练参数将mPLUG-video扩展为中文多模态大语言模型,展示了其强大的指令理解与视频语义分析能力。零样本指令理解实验表明,基于Youku-mPLUG的预训练可增强模型对整体与细粒度视觉语义的理解、场景文本识别及开放域知识利用能力。