Simultaneous speech translation (SimulST) translates partial speech inputs incrementally. Although the monotonic correspondence between input and output is preferable for smaller latency, it is not the case for distant language pairs such as English and Japanese. A prospective approach to this problem is to mimic simultaneous interpretation (SI) using SI data to train a SimulST model. However, the size of such SI data is limited, so the SI data should be used together with ordinary bilingual data whose translations are given in offline. In this paper, we propose an effective way to train a SimulST model using mixed data of SI and offline. The proposed method trains a single model using the mixed data with style tags that tell the model to generate SI- or offline-style outputs. Experiment results show improvements of BLEURT in different latency ranges, and our analyses revealed the proposed model generates SI-style outputs more than the baseline.
翻译:同声传译(SimulST)需逐步翻译部分语音输入。尽管输入与输出之间的单调对应有利于降低延迟,但对于英语与日语等距离较远的语言对而言,这一要求难以实现。针对该问题的潜在方法是利用同声传译(SI)数据训练模拟同声传译的SimulST模型。然而,此类SI数据规模有限,因此需将SI数据与提供离线翻译的普通双语数据结合使用。本文提出一种利用SI与离线混合数据训练SimulST模型的有效方法。该方法通过带风格标签的混合数据训练单一模型,指示模型生成SI风格或离线风格的输出。实验结果表明,该方法在不同延迟区间内均提升了BLEURT评分,且分析显示所提模型比基线模型生成了更多SI风格的输出。