BERTScore is an effective and robust automatic metric for referencebased machine translation evaluation. In this paper, we incorporate multilingual knowledge graph into BERTScore and propose a metric named KG-BERTScore, which linearly combines the results of BERTScore and bilingual named entity matching for reference-free machine translation evaluation. From the experimental results on WMT19 QE as a metric without references shared tasks, our metric KG-BERTScore gets higher overall correlation with human judgements than the current state-of-the-art metrics for reference-free machine translation evaluation.1 Moreover, the pre-trained multilingual model used by KG-BERTScore and the parameter for linear combination are also studied in this paper.
翻译:BERTScore是一种有效且鲁棒的无参考机器翻译评估自动指标。本文通过将多语言知识图谱融入BERTScore,提出了一种名为KG-BERTScore的评估指标,该指标通过线性组合BERTScore结果与双语命名实体匹配,实现无参考机器翻译评估。基于WMT19质量评估作为无参考指标共享任务的实验结果表明,KG-BERTScore与人工判断的整体相关性优于当前最先进的无参考机器翻译评估指标。此外,本文还研究了KG-BERTScore使用的预训练多语言模型及线性组合参数。