Many users reading online articles in various magazines may suffer considerable difficulty in distinguishing the implicit intents in texts. In this work, we focus on automatically recognizing the political intents of a given online newspaper by understanding the context of the text. To solve this task, we present a novel Korean text classification dataset that contains various articles. We also provide deep-learning-based text classification baseline models trained on the proposed dataset. Our dataset contains 12,000 news articles that may contain political intentions, from the politics section of six of the most representative newspaper organizations in South Korea. All the text samples are labeled simultaneously in two aspects (1) the level of political orientation and (2) the level of pro-government. To the best of our knowledge, our paper is the most large-scale Korean news dataset that contains long text and addresses multi-task classification problems. We also train recent state-of-the-art (SOTA) language models that are based on transformer architectures and demonstrate that the trained models show decent text classification performance. All the codes, datasets, and trained models are available at https://github.com/Kdavid2355/KoPolitic-Benchmark-Dataset.
翻译:许多在线阅读各类杂志文章的用户可能难以区分文本中的隐含意图。本研究聚焦于通过理解文本上下文自动识别给定网络报纸的政治意图。为解决该任务,我们提出了一个包含多种文章的新型韩语文本分类数据集,并提供了基于该数据集训练的深度学习文本分类基线模型。本数据集包含来自韩国六家最具代表性报社政治板块的12,000篇可能蕴含政治意图的新闻文章。所有文本样本均从两个维度进行同步标注:(1)政治倾向程度;(2)亲政府程度。据我们所知,本文提出的数据集是规模最大的韩语长文本新闻数据集,同时解决了多任务分类问题。我们还训练了基于Transformer架构的最新最先进语言模型,实验表明这些模型展现了优秀的文本分类性能。所有代码、数据集和训练模型均发布在https://github.com/Kdavid2355/KoPolitic-Benchmark-Dataset。