This paper explains the participation of team Hitachi to SemEval-2023 Task 3 "Detecting the genre, the framing, and the persuasion techniques in online news in a multi-lingual setup.'' Based on the multilingual, multi-task nature of the task and the low-resource setting, we investigated different cross-lingual and multi-task strategies for training the pretrained language models. Through extensive experiments, we found that (a) cross-lingual/multi-task training, and (b) collecting an external balanced dataset, can benefit the genre and framing detection. We constructed ensemble models from the results and achieved the highest macro-averaged F1 scores in Italian and Russian genre categorization subtasks.
翻译:本论文阐述了日立团队参与SemEval-2023任务3"多语言环境下在线新闻体裁、框架及说服技巧检测"的相关工作。基于该任务的多语言、多任务特性以及低资源场景,我们研究了训练预训练语言模型的不同跨语言与多任务策略。通过广泛实验发现:(a)跨语言/多任务训练,以及(b)收集外部平衡数据集,均有利于体裁与框架检测任务。我们基于实验结果构建集成模型,在意大利语和俄语体裁分类子任务中获得了最高的宏平均F1分数。