We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this dataset to investigate whether machine translation (MT) metrics which are fine-tuned on human-generated MT quality judgements are robust to domain shifts between training and inference. We find that fine-tuned metrics exhibit a substantial performance drop in the unseen domain scenario relative to metrics that rely on the surface form, as well as pre-trained metrics which are not fine-tuned on MT quality judgments.
翻译:我们引入了一个新的、涵盖11个语言对生物医学领域的大规模多维质量指标(MQM)标注数据集。利用该数据集,我们探究了基于人工生成机器翻译质量评判进行微调的机器翻译(MT)指标,在训练与推理之间存在领域迁移时是否具有鲁棒性。研究发现,与依赖表层形式的指标以及未经MT质量评判微调的预训练指标相比,微调后的指标在未见领域场景下性能显著下降。