Ideological divisions in the United States have become increasingly prominent in daily communication. Accordingly, there has been much research on political polarization, including many recent efforts that take a computational perspective. By detecting political biases in a corpus of text, one can attempt to describe and discern the polarity of that text. Intuitively, the named entities (i.e., the nouns and the phrases that act as nouns) and hashtags in text often carry information about political views. For example, people who use the term "pro-choice" are likely to be liberal, whereas people who use the term "pro-life" are likely to be conservative. In this paper, we seek to reveal political polarities in social-media text data and to quantify these polarities by explicitly assigning a polarity score to entities and hashtags. Although this idea is straightforward, it is difficult to perform such inference in a trustworthy quantitative way. Key challenges include the small number of known labels, the continuous spectrum of political views, and the preservation of both a polarity score and a polarity-neutral semantic meaning in an embedding vector of words. To attempt to overcome these challenges, we propose the Polarity-aware Embedding Multi-task learning (PEM) model. This model consists of (1) a self-supervised context-preservation task, (2) an attention-based tweet-level polarity-inference task, and (3) an adversarial learning task that promotes independence between an embedding's polarity dimension and its semantic dimensions. Our experimental results demonstrate that our PEM model can successfully learn polarity-aware embeddings that perform well classification tasks. We examine a variety of applications and we thereby demonstrate the effectiveness of our PEM model. We also discuss important limitations of our work and encourage caution when applying the it to real-world scenarios.
翻译:美国社会中的意识形态分歧在日常交流中日益凸显。因此,围绕政治极化的研究大量涌现,其中许多近期工作采用计算视角。通过检测文本语料中的政治偏见,我们可以尝试描述和识别文本的政治倾向。直观来看,文本中的命名实体(即充当名词的名词和短语)和话题标签往往携带政治观点信息。例如,使用“pro-choice”一词的人可能倾向自由派,而使用“pro-life”的人则可能倾向保守派。本文旨在揭示社交媒体文本数据中的政治极化,并通过为实体和话题标签明确分配极化分数来量化这些极化。尽管这一思路直接明了,但如何以可信的量化方式进行推断却颇具挑战。关键挑战包括已知标签数量有限、政治观点的连续谱系,以及在词嵌入向量中同时保留极化分数和与极性无关的语义信息。为克服这些挑战,我们提出极化感知嵌入多任务学习(PEM)模型。该模型包含:(1)自监督的语境保留任务,(2)基于注意力的推文级极化推断任务,以及(3)促进嵌入的极化维度与语义维度相互独立的对抗学习任务。实验结果表明,我们的PEM模型能够成功学习到极化感知嵌入,并在分类任务中表现优异。我们检验了多种应用场景,从而证明了PEM模型的有效性。同时,我们讨论了本研究的关键局限性,并提醒在将其应用于现实场景时需保持谨慎。