With the widespread use of social networks, detecting the topics discussed in these networks has become a significant challenge. The current works are mainly based on frequent pattern mining or semantic relations, and the language structure is not considered. The meaning of language structural methods is to discover the relationship between words and how humans understand them. Therefore, this paper uses the Concept of the Imitation of the Mental Ability of Word Association to propose a topic detection framework in social networks. This framework is based on the Human Word Association method. A special extraction algorithm has also been designed for this purpose. The performance of this method is evaluated on the FA-CUP dataset. It is a benchmark dataset in the field of topic detection. The results show that the proposed method is a good improvement compared to other methods, based on the Topic-recall and the keyword F1 measure. Also, most of the previous works in the field of topic detection are limited to the English language, and the Persian language, especially microblogs written in this language, is considered a low-resource language. Therefore, a data set of Telegram posts in the Farsi language has been collected. Applying the proposed method to this dataset also shows that this method works better than other topic detection methods.
翻译:随着社交网络的广泛使用,检测这些网络中讨论的话题已成为一项重要挑战。现有研究主要基于频繁模式挖掘或语义关系,未考虑语言结构。语言结构方法的意义在于发现词汇之间的关系以及人类如何理解它们。因此,本文借鉴"词汇关联心智能力模仿"概念,提出一种社交网络话题检测框架。该框架基于人类词汇关联方法,并为此设计了专门的提取算法。在话题检测领域的基准数据集FA-CUP上评估了该方法的表现。结果表明,基于话题召回率和关键词F1值,所提方法相比其他方法有显著改进。此外,以往话题检测研究大多局限于英语,而波斯语(尤其是该语言编写的微博内容)被视为低资源语言。为此,我们收集了波斯语Telegram帖子数据集,将该方法应用于此数据集的结果显示,该方法优于其他话题检测方法。