Using crowdsourcing, we collected more than 10,000 URL pairs (parallel top page pairs) of bilingual websites that contain parallel documents and created a Japanese-Chinese parallel corpus of 4.6M sentence pairs from these websites. We used a Japanese-Chinese bilingual dictionary of 160K word pairs for document and sentence alignment. We then used high-quality 1.2M Japanese-Chinese sentence pairs to train a parallel corpus filter based on statistical language models and word translation probabilities. We compared the translation accuracy of the model trained on these 4.6M sentence pairs with that of the model trained on Japanese-Chinese sentence pairs from CCMatrix (12.4M), a parallel corpus from global web mining. Although our corpus is only one-third the size of CCMatrix, we found that the accuracy of the two models was comparable and confirmed that it is feasible to use crowdsourcing for web mining of parallel data.
翻译:我们采用众包方法,收集了超过10,000组包含平行文档的双语网站URL对(平行首页对),并基于这些网站构建了包含460万句对的日汉平行语料库。我们使用包含16万词对的日汉双语词典进行文档和句级对齐,随后利用120万高质量日汉句对,训练了基于统计语言模型和词翻译概率的平行语料过滤器。将基于这460万句对训练的模型与基于全球网络挖掘的平行语料库CCMatrix(1240万句对)的日汉句对训练模型进行翻译精度对比,发现尽管本语料库规模仅为CCMatrix的三分之一,但两个模型的翻译精度相当,证实了利用众包进行网络挖掘获取平行数据的可行性。