Recent advances in machine translation (MT) have shown that Minimum Bayes Risk (MBR) decoding can be a powerful alternative to beam search decoding, especially when combined with neural-based utility functions. However, the performance of MBR decoding depends heavily on how and how many candidates are sampled from the model. In this paper, we explore how different sampling approaches for generating candidate lists for MBR decoding affect performance. We evaluate popular sampling approaches, such as ancestral, nucleus, and top-k sampling. Based on our insights into their limitations, we experiment with the recently proposed epsilon-sampling approach, which prunes away all tokens with a probability smaller than epsilon, ensuring that each token in a sample receives a fair probability mass. Through extensive human evaluations, we demonstrate that MBR decoding based on epsilon-sampling significantly outperforms not only beam search decoding, but also MBR decoding with all other tested sampling methods across four language pairs.
翻译:近期机器翻译(MT)的进展表明,最小贝叶斯风险(MBR)解码可作为波束搜索解码的有力替代方案,尤其在与基于神经网络的效用函数结合时表现突出。然而,MBR 解码的性能高度依赖于候选样本的生成方式与数量。本文探讨了不同采样策略对 MBR 解码候选列表生成效果的影响。我们评估了多种主流采样方法,包括祖先采样、核采样和 top-k 采样。基于对其局限性的分析,我们实验了近期提出的 epsilon 采样方法——该方法剔除所有概率小于 epsilon 的标记,从而确保采样中每个标记获得公平的概率质量。通过大规模人工评估,我们证明:基于 epsilon 采样的 MBR 解码不仅显著优于波束搜索解码,而且在四个语言对上的表现均超过所有其他测试采样方法的 MBR 解码。