Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges -- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work.
翻译:最近,语音到文本翻译引起了越来越多的关注,相关研究迅速涌现。本文对直接语音翻译进行了全面综述,旨在总结当前最先进的技术。首先,我们根据主要挑战将现有研究工作分为三类——建模负担、数据稀缺和应用问题。针对建模负担问题,提出了两种主要结构:编码器-解码器框架(Transformer及其变体)和多任务框架。针对数据稀缺的挑战,最近的研究采用了多种复杂技术,例如数据增强、预训练、知识蒸馏和多语言建模。我们分析并总结了应用问题,包括实时性、分割、命名实体、性别偏见和代码切换。最后,我们讨论了未来一些有前景的研究方向。