Topic Modelling (TM) is from the research branches of natural language understanding (NLU) and natural language processing (NLP) that is to facilitate insightful analysis from large documents and datasets, such as a summarisation of main topics and the topic changes. This kind of discovery is getting more popular in real-life applications due to its impact on big data analytics. In this study, from the social-media and healthcare domain, we apply popular Latent Dirichlet Allocation (LDA) methods to model the topic changes in Swedish newspaper articles about Coronavirus. We describe the corpus we created including 6515 articles, methods applied, and statistics on topic changes over approximately 1 year and two months period of time from 17th January 2020 to 13th March 2021. We hope this work can be an asset for grounding applications of topic modelling and can be inspiring for similar case studies in an era with pandemics, to support socio-economic impact research as well as clinical and healthcare analytics. Our data and source code are openly available at https://github. com/poethan/Swed_Covid_TM Keywords: Latent Dirichlet Allocation (LDA); Topic Modelling; Coronavirus; Pandemics; Natural Language Understanding; BERT-topic
翻译:主题建模(Topic Modelling, TM)是自然语言理解(NLU)和自然语言处理(NLP)研究领域的分支,旨在促进对大规模文档和数据集的深入分析,例如主要主题的总结及主题变化。由于其对大数据分析的影响,这种发现方法在实际应用中日渐流行。在本研究中,我们聚焦社交媒体与医疗健康领域,应用广泛使用的潜在狄利克雷分配(Latent Dirichlet Allocation, LDA)方法,对瑞典语新冠病毒相关报纸文章的主题变化进行建模。我们描述了所构建的包含6515篇文章的语料库、应用的方法,以及从2020年1月17日至2021年3月13日约一年零两个月时间内的主题变化统计。我们希望这项工作能为主题建模的实际应用提供基础,并在疫情时代对类似案例研究起到启发作用,以支持社会经济影响研究以及临床和医疗健康分析。我们的数据与源代码已公开在 https://github.com/poethan/Swed_Covid_TM。关键词:潜在狄利克雷分配(LDA);主题建模;新冠病毒;疫情;自然语言理解;BERT-topic