This paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP-LiAuto (Team 237) in the first Chinese Continuous Visual Speech Recognition Challenge (CNVSRC) 2023, engaging in the fixed and open tracks of Single-Speaker VSR Task, and the open track of Multi-Speaker VSR Task. In terms of data processing, we leverage the lip motion extractor from the baseline1 to produce multi-scale video data. Besides, various augmentation techniques are applied during training, encompassing speed perturbation, random rotation, horizontal flipping, and color transformation. The VSR model adopts an end-to-end architecture with joint CTC/attention loss, comprising a ResNet3D visual frontend, an E-Branchformer encoder, and a Transformer decoder. Experiments show that our system achieves 34.76% CER for the Single-Speaker Task and 41.06% CER for the Multi-Speaker Task after multi-system fusion, ranking first place in all three tracks we participate.
翻译:本文阐述了NPU-ASLP-LiAuto(第237组)在首届中文连续视觉语音识别挑战赛(CNVSRC 2023)中提出的视觉语音识别(VSR)系统,该系统参与了单说话人VSR任务的固定轨道与开放轨道,以及多说话人VSR任务的开放轨道。在数据处理方面,我们利用基线系统的唇部运动提取器生成多尺度视频数据。此外,在训练过程中应用了多种数据增强技术,包括速度扰动、随机旋转、水平翻转和色彩变换。VSR模型采用基于CTC/注意力联合损失的端到端架构,包含ResNet3D视觉前端、E-Branchformer编码器和Transformer解码器。实验表明,经多系统融合后,我们的系统在单说话人任务中字错误率(CER)为34.76%,多说话人任务中为41.06%,在参与的所有三个轨道中均排名第一。