Sign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes the main bottleneck for SLR. Most SLR works thereby adopt pretrained visual modules and develop two mainstream solutions. The multi-stream architectures extend multi-cue visual features, yielding the current SOTA performances but requiring complex designs and might introduce potential noise. Alternatively, the advanced single-cue SLR frameworks using explicit cross-modal alignment between visual and textual modalities are simple and effective, potentially competitive with the multi-cue framework. In this work, we propose a novel contrastive visual-textual transformation for SLR, CVT-SLR, to fully explore the pretrained knowledge of both the visual and language modalities. Based on the single-cue cross-modal alignment framework, we propose a variational autoencoder (VAE) for pretrained contextual knowledge while introducing the complete pretrained language module. The VAE implicitly aligns visual and textual modalities while benefiting from pretrained contextual knowledge as the traditional contextual module. Meanwhile, a contrastive cross-modal alignment algorithm is designed to explicitly enhance the consistency constraints. Extensive experiments on public datasets (PHOENIX-2014 and PHOENIX-2014T) demonstrate that our proposed CVT-SLR consistently outperforms existing single-cue methods and even outperforms SOTA multi-cue methods.
翻译:手语识别(SLR)是一项弱监督任务,旨在将手语视频标注为文本注释。近期研究表明,缺乏大规模可用手语数据集导致的训练不足已成为SLR的主要瓶颈。因此,多数SLR工作采用预训练视觉模块,并发展出两种主流方案:多流架构通过扩展多线索视觉特征取得了当前最优性能,但需复杂设计且可能引入潜在噪声;而采用显式跨模态对齐的先进单线索SLR框架简洁高效,具有与多线索框架竞争的潜力。本文提出一种新颖的对比式视觉-文本转换方法CVT-SLR,以充分挖掘视觉与语言模态的预训练知识。基于单线索跨模态对齐框架,我们引入变分自编码器(VAE)处理预训练上下文知识,同时采用完整的预训练语言模块。该VAE在隐式对齐视觉与文本模态的同时,能像传统上下文模块一样受益于预训练知识。此外,我们设计了对比式跨模态对齐算法以显式增强一致性约束。在公开数据集(PHOENIX-2014和PHOENIX-2014T)上的大量实验表明,所提CVT-SLR始终优于现有单线索方法,甚至超越了最先进的多线索方法。