Context-aware translation can be achieved by processing a concatenation of consecutive sentences with the standard translation approach. This paper investigates the intuitive idea of adopting segment embeddings for this task to help the Transformer discern the position of each sentence in the concatenation sequence. We compare various segment embeddings and propose novel methods to encode sentence position into token representations, showing that they do not benefit the vanilla concatenation approach except in a specific setting.
翻译:上下文感知翻译可通过采用标准翻译方法处理连续句子的拼接来实现。本文探究了在此任务中采用分段嵌入以帮助Transformer区分拼接序列中各句子位置的直观思路。我们比较了多种分段嵌入方法,并提出了将句子位置编码到词元表示中的新方法,结果表明除特定设置外,这些方法对基础拼接方法并无增益。