We introduce a dataset comprising commercial machine translations, gathered weekly over six years across 12 translation directions. Since human A/B testing is commonly used, we assume commercial systems improve over time, which enables us to evaluate machine translation (MT) metrics based on their preference for more recent translations. Our study confirms several previous findings in MT metrics research and demonstrates the dataset's value as a testbed for metric evaluation. We release our code at https://github.com/gjwubyron/Evo
翻译:我们引入了一个包含商业机器翻译的数据集,该数据集在六年内每周收集,涵盖12个翻译方向。由于人工A/B测试被广泛使用,我们假设商业系统会随时间改进,这使我们能够基于指标对近期翻译的偏好来评估机器翻译(MT)指标。我们的研究证实了先前MT指标研究中的若干发现,并展示了该数据集作为指标评估测试平台的价值。我们在https://github.com/gjwubyron/Evo发布了代码。