Hand gesture serves as a crucial role during the expression of sign language. Current deep learning based methods for sign language understanding (SLU) are prone to over-fitting due to insufficient sign data resource and suffer limited interpretability. In this paper, we propose the first self-supervised pre-trainable SignBERT+ framework with model-aware hand prior incorporated. In our framework, the hand pose is regarded as a visual token, which is derived from an off-the-shelf detector. Each visual token is embedded with gesture state and spatial-temporal position encoding. To take full advantage of current sign data resource, we first perform self-supervised learning to model its statistics. To this end, we design multi-level masked modeling strategies (joint, frame and clip) to mimic common failure detection cases. Jointly with these masked modeling strategies, we incorporate model-aware hand prior to better capture hierarchical context over the sequence. After the pre-training, we carefully design simple yet effective prediction heads for downstream tasks. To validate the effectiveness of our framework, we perform extensive experiments on three main SLU tasks, involving isolated and continuous sign language recognition (SLR), and sign language translation (SLT). Experimental results demonstrate the effectiveness of our method, achieving new state-of-the-art performance with a notable gain.
翻译:手势在表达手语中起着关键作用。当前基于深度学习的符号语言理解方法由于手语数据资源不足,容易过拟合且可解释性有限。本文提出首个具备模型感知手部先验的自监督预训练框架SignBERT+。在该框架中,手势被视为视觉标记,由现成的检测器提取。每个视觉标记嵌入手势状态和时空位置编码。为充分利用当前手语数据资源,我们首先进行自监督学习以建模其统计特性。为此,设计多层级掩码建模策略(关节、帧和片段)以模拟常见检测失败情况。结合这些掩码建模策略,我们引入模型感知手部先验,以更好地捕捉序列上的层级上下文。预训练后,我们精心设计简单而有效的预测头用于下游任务。为验证框架有效性,在三个主要手语理解任务上进行了大量实验,包括孤立词和连续手语识别,以及手语翻译。实验结果表明了该方法的有效性,以显著优势实现了新的最先进性能。