Sign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes the main bottleneck for SLR. Most SLR works thereby adopt pretrained visual modules and develop two mainstream solutions. The multi-stream architectures extend multi-cue visual features, yielding the current SOTA performances but requiring complex designs and might introduce potential noise. Alternatively, the advanced single-cue SLR frameworks using explicit cross-modal alignment between visual and textual modalities are simple and effective, potentially competitive with the multi-cue framework. In this work, we propose a novel contrastive visual-textual transformation for SLR, CVT-SLR, to fully explore the pretrained knowledge of both the visual and language modalities. Based on the single-cue cross-modal alignment framework, we propose a variational autoencoder (VAE) for pretrained contextual knowledge while introducing the complete pretrained language module. The VAE implicitly aligns visual and textual modalities while benefiting from pretrained contextual knowledge as the traditional contextual module. Meanwhile, a contrastive cross-modal alignment algorithm is designed to explicitly enhance the consistency constraints. Extensive experiments on public datasets (PHOENIX-2014 and PHOENIX-2014T) demonstrate that our proposed CVT-SLR consistently outperforms existing single-cue methods and even outperforms SOTA multi-cue methods.
翻译:手语识别(SLR)是一项弱监督任务,旨在将手语视频标注为文本注释。近期研究表明,缺乏大规模可用手语数据集导致的训练不足成为SLR的主要瓶颈。为此,多数SLR工作采用预训练视觉模块,并发展出两种主流解决方案。多流架构通过扩展多线索视觉特征取得当前最优性能,但需复杂设计且可能引入潜在噪声。相比之下,基于视觉与文本模态间显式跨模态对齐的先进单线索SLR框架兼具简洁性与有效性,具备与多线索框架竞争的能力。本文提出一种面向SLR的新型对比视觉-文本变换方法CVT-SLR,以充分挖掘视觉与语言模态的预训练知识。基于单线索跨模态对齐框架,我们引入变分自编码器(VAE)以利用预训练上下文知识,同时引入完整预训练语言模块。该VAE在隐式对齐视觉与文本模态的同时,继承传统上下文模块所依赖的预训练知识优势。此外,我们设计了一种对比跨模态对齐算法,以显式增强一致性约束。在公开数据集(PHOENIX-2014和PHOENIX-2014T)上的大量实验表明,所提出的CVT-SLR始终优于现有单线索方法,甚至显著超越当前最优的多线索方法。