Automatic translation from signed to spoken languages is an interdisciplinary research domain, lying on the intersection of computer vision, machine translation and linguistics. Nevertheless, research in this domain is performed mostly by computer scientists in isolation. As the domain is becoming increasingly popular - the majority of scientific papers on the topic of sign language translation have been published in the past three years - we provide an overview of the state of the art as well as some required background in the different related disciplines. We give a high-level introduction to sign language linguistics and machine translation to illustrate the requirements of automatic sign language translation. We present a systematic literature review to illustrate the state of the art in the domain and then, harking back to the requirements, lay out several challenges for future research. We find that significant advances have been made on the shoulders of spoken language machine translation research. However, current approaches are often not linguistically motivated or are not adapted to the different input modality of sign languages. We explore challenges related to the representation of sign language data, the collection of datasets, the need for interdisciplinary research and requirements for moving beyond research, towards applications. Based on our findings, we advocate for interdisciplinary research and to base future research on linguistic analysis of sign languages. Furthermore, the inclusion of deaf and hearing end users of sign language translation applications in use case identification, data collection and evaluation is of the utmost importance in the creation of useful sign language translation models. We recommend iterative, human-in-the-loop, design and development of sign language translation models.
翻译:手语到有声语言的自动翻译是一个跨学科研究领域,涉及计算机视觉、机器翻译和语言学的交叉。然而,该领域的研究主要由计算机科学家独立进行。随着该领域日益普及——过去三年中发表的大多数关于手语翻译主题的科学论文——我们概述了当前技术水平以及不同相关学科的一些必要背景。我们对手语语言学和机器翻译进行了高层次介绍,以说明自动手语翻译的要求。我们进行了系统文献综述,以展示该领域的现状,然后回顾这些要求,提出了未来研究的若干挑战。我们发现,在口语机器翻译研究的基础上取得了显著进展。然而,当前的方法通常缺乏语言学动机,或者未适应手语的不同输入模态。我们探讨了与手语数据表示、数据集收集、跨学科研究需求以及从研究迈向应用所需的更高要求相关的挑战。基于我们的发现,我们主张进行跨学科研究,并将未来研究建立在对手语的语言学分析之上。此外,在用例识别、数据收集和评估中纳入手语翻译应用的手语使用者和听力健全最终用户,对于创建有用的手语翻译模型至关重要。我们建议采用迭代式、人在回路中的设计和开发方法构建手语翻译模型。