Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word embeddings are complex, high-dimensional spaces and a focus on identifying differences only captures a fraction of their richness. Here, we take a step towards leveraging the richness of the full embedding space, by using word embeddings to map out how words are used differently. Specifically, we describe the construction of dialectograms, an unsupervised way to visually explore the characteristic ways in which each community use a focal word. Based on these dialectograms, we provide a new measure of the degree to which words are used differently that overcomes the tendency for existing measures to pick out low frequent or polysemous words. We apply our methods to explore the discourses of two US political subreddits and show how our methods identify stark affective polarisation of politicians and political entities, differences in the assessment of proper political action as well as disagreement about whether certain issues require political intervention at all.
翻译:词嵌入提供了一种无监督的方式,用以理解话语社群之间用词差异。近期多项研究聚焦于识别两个或更多社群中存在不同用法的词汇。然而,词嵌入是高维复杂空间,仅关注差异识别只能捕捉其丰富性的冰山一角。本文旨在通过词嵌入映射词汇用法的差异模式,探索完整嵌入空间的丰富性。具体而言,我们描述了方言图(dialectograms)的构建方法——这是一种无监督的可视化工具,用于展示各社群使用核心词汇的典型特征。基于这些方言图,我们提出了一种度量词汇用法差异程度的新方法,有效克服了现有方法偏向识别低频词或多义词的倾向。我们将该方法应用于分析美国两个政治子论坛(subreddits)的话语,结果表明该方法能清晰揭示政治人物与政治实体的鲜明情感极化,展现对正当政治行动评估的分歧,以及关于某些议题是否需要政治干预的认知差异。