Capturing and annotating Sign language datasets is a time consuming and costly process. Current datasets are orders of magnitude too small to successfully train unconstrained \acf{slt} models. As a result, research has turned to TV broadcast content as a source of large-scale training data, consisting of both the sign language interpreter and the associated audio subtitle. However, lack of sign language annotation limits the usability of this data and has led to the development of automatic annotation techniques such as sign spotting. These spottings are aligned to the video rather than the subtitle, which often results in a misalignment between the subtitle and spotted signs. In this paper we propose a method for aligning spottings with their corresponding subtitles using large spoken language models. Using a single modality means our method is computationally inexpensive and can be utilized in conjunction with existing alignment techniques. We quantitatively demonstrate the effectiveness of our method on the \acf{mdgs} and \acf{bobsl} datasets, recovering up to a 33.22 BLEU-1 score in word alignment.
翻译:捕捉和标注手语数据集是一个耗时且成本高昂的过程。当前数据集的规模小了几个数量级,无法成功训练无约束的手语翻译(SLT)模型。因此,研究转向电视广播内容作为大规模训练数据的来源,其中包含手语翻译员和相关音频字幕。然而,手语标注的缺乏限制了这些数据的可用性,并促使了自动标注技术的发展,例如手语检测。这些检测是与视频对齐,而非与字幕对齐,这通常导致字幕与检测到的手语之间出现对齐偏差。在本文中,我们提出了一种方法,利用大型口语语言模型将检测结果与对应的字幕进行对齐。使用单一模态意味着我们的方法计算成本低廉,并且可以与现有的对齐技术结合使用。我们通过在多粒度手语(MDGS)和Bohmer手语(BOBSL)数据集上的实验,定量证明了我们方法的有效性,在词汇对齐中恢复了高达33.22的BLEU-1分数。