It is important to be able to analyze the emotional state of people around the globe. There are 7100+ active languages spoken around the world and building emotion classification for each language is labor intensive. Particularly for low-resource and endangered languages, building emotion classification can be quite challenging. We present a cross-lingual emotion classifier, where we train an emotion classifier with resource-rich languages (i.e. \textit{English} in our work) and transfer the learning to low and moderate resource languages. We compare and contrast two approaches of transfer learning from a high-resource language to a low or moderate-resource language. One approach projects the annotation from a high-resource language to low and moderate-resource language in parallel corpora and the other one uses direct transfer from high-resource language to the other languages. We show the efficacy of our approaches on 6 languages: Farsi, Arabic, Spanish, Ilocano, Odia, and Azerbaijani. Our results indicate that our approaches outperform random baselines and transfer emotions across languages successfully. For all languages, the direct cross-lingual transfer of emotion yields better results. We also create annotated emotion-labeled resources for four languages: Farsi, Azerbaijani, Ilocano and Odia.
翻译:能够分析全球人群的情感状态至关重要。全球现存7100多种活跃语言,为每种语言构建情感分类系统需要大量人力投入,尤其对于低资源和濒危语言而言,构建情感分类系统极具挑战性。本文提出一种跨语言情感分类器:我们利用资源丰富语言(本文以英语为例)训练情感分类模型,并将学习能力迁移至低资源和中资源语言。我们比较了两种从高资源语言向低/中资源语言迁移学习的方法:一种通过平行语料将高资源语言的标注映射到低/中资源语言,另一种直接进行跨语言迁移。我们在六种语言(波斯语、阿拉伯语、西班牙语、伊洛卡诺语、奥里亚语、阿塞拜疆语)上验证了方法的有效性。结果表明,我们的方法优于随机基线,并能成功实现跨语言情感迁移。对于所有目标语言,直接跨语言情感迁移的效果更优。我们还为波斯语、阿塞拜疆语、伊洛卡诺语和奥里亚语四种语言创建了带情感标注的资源数据集。