The ultimate goal of continuous sign language recognition(CSLR) is to facilitate the communication between special people and normal people, which requires a certain degree of real-time and deploy-ability of the model. However, in the previous research on CSLR, little attention has been paid to the real-time and deploy-ability. In order to improve the real-time and deploy-ability of the model, this paper proposes a zero parameter, zero computation temporal superposition crossover module(TSCM), and combines it with 2D convolution to form a "TSCM+2D convolution" hybrid convolution, which enables 2D convolution to have strong spatial-temporal modelling capability with zero parameter increase and lower deployment cost compared with other spatial-temporal convolutions. The overall CSLR model based on TSCM is built on the improved ResBlockT network in this paper. The hybrid convolution of "TSCM+2D convolution" is applied to the ResBlock of the ResNet network to form the new ResBlockT, and random gradient stop and multi-level CTC loss are introduced to train the model, which reduces the final recognition WER while reducing the training memory usage, and extends the ResNet network from image classification task to video recognition task. In addition, this study is the first in CSLR to use only 2D convolution extraction of sign language video temporal-spatial features for end-to-end learning for recognition. Experiments on two large-scale continuous sign language datasets demonstrate the effectiveness of the proposed method and achieve highly competitive results.
翻译:连续手语识别(CSLR)的最终目标是促进特殊人群与普通人群之间的交流,这要求模型具备一定程度的实时性与可部署性。然而,在以往的CSLR研究中,鲜有关注模型的实时性与可部署性。为提升模型的实时性与可部署性,本文提出一种零参数、零计算的时序叠加交叉模块(Temporal Superposition Crossover Module, TSCM),并将其与二维卷积结合形成“TSCM+二维卷积”混合卷积结构。该结构使二维卷积在零参数增加且部署成本低于其他时空卷积的条件下,具备强大的时空建模能力。本文基于改进的ResBlockT网络构建了基于TSCM的完整CSLR模型。将“TSCM+二维卷积”混合卷积应用于ResNet网络的ResBlock中形成新型ResBlockT,并引入随机梯度停止与多层级CTC损失函数进行模型训练。该方法在降低训练内存占用的同时减小了最终识别的词错误率(WER),并将ResNet网络从图像分类任务拓展至视频识别任务。此外,本研究首次在CSLR中仅使用二维卷积提取手语视频时空特征,实现端到端识别学习。在两个大规模连续手语数据集上的实验验证了所提方法的有效性,并取得了极具竞争力的结果。