Generating grammatically and semantically correct captions in video captioning is a challenging task. The captions generated from the existing methods are either word-by-word that do not align with grammatical structure or miss key information from the input videos. To address these issues, we introduce a novel global-local fusion network, with a Global-Local Fusion Block (GLFB) that encodes and fuses features from different parts of speech (POS) components with visual-spatial features. We use novel combinations of different POS components - 'determinant + subject', 'auxiliary verb', 'verb', and 'determinant + object' for supervision of the POS blocks - Det + Subject, Aux Verb, Verb, and Det + Object respectively. The novel global-local fusion network together with POS blocks helps align the visual features with language description to generate grammatically and semantically correct captions. Extensive qualitative and quantitative experiments on benchmark MSVD and MSRVTT datasets demonstrate that the proposed approach generates more grammatically and semantically correct captions compared to the existing methods, achieving the new state-of-the-art. Ablations on the POS blocks and the GLFB demonstrate the impact of the contributions on the proposed method.
翻译:在视频描述生成中生成语法与语义正确的描述是一项具有挑战性的任务。现有方法生成的描述要么逐词生成且不符合语法结构,要么遗漏输入视频中的关键信息。为解决这些问题,我们提出了一种新颖的全局-局部融合网络,该网络包含一个全局-局部融合块(GLFB),用于编码不同词性(POS)组件的特征并将其与视觉-空间特征相融合。我们创新性地组合了不同的词性组件——"限定词+主语"、"助动词"、"动词"和"限定词+宾语"——分别用于监督对应的词性模块:Det+Subject、Aux Verb、Verb 以及 Det+Object。这种新颖的全局-局部融合网络与词性模块协同作用,有助于将视觉特征与语言描述对齐,从而生成语法与语义正确的描述。在基准数据集MSVD和MSRVTT上进行的大量定性与定量实验表明,与现有方法相比,所提方法能生成更多语法与语义正确的描述,达到了新的最优性能。针对词性模块和GLFB的消融实验证明了各贡献对所提方法的影响。