Sign languages are the primary means of communication for many hard-of-hearing people worldwide. Recently, to bridge the communication gap between the hard-of-hearing community and the rest of the population, several sign language translation datasets have been proposed to enable the development of statistical sign language translation systems. However, there is a dearth of sign language resources for the Indian sign language. This resource paper introduces ISLTranslate, a translation dataset for continuous Indian Sign Language (ISL) consisting of 31k ISL-English sentence/phrase pairs. To the best of our knowledge, it is the largest translation dataset for continuous Indian Sign Language. We provide a detailed analysis of the dataset. To validate the performance of existing end-to-end Sign language to spoken language translation systems, we benchmark the created dataset with a transformer-based model for ISL translation.
翻译:手语是全球许多听障人士的主要交流方式。近年来,为缩小听障群体与普通人群之间的沟通鸿沟,研究者提出了多个手语翻译数据集,以支持统计手语翻译系统的开发。然而,针对印度手语的语料资源仍然匮乏。本资源论文介绍了ISLTranslate——一个面向连续印度手语(ISL)的翻译数据集,包含31,000个ISL-英语句子/短语对。据我们所知,这是目前规模最大的连续印度手语翻译数据集。我们对数据集进行了详细分析。为验证现有端到端手语到口语翻译系统的性能,我们基于该数据集,采用Transformer模型对ISL翻译任务进行了基准测试。