With the advent and popularity of big data mining and huge text analysis in modern times, automated text summarization became prominent for extracting and retrieving important information from documents. This research investigates aspects of automatic text summarization from the perspectives of single and multiple documents. Summarization is a task of condensing huge text articles into short, summarized versions. The text is reduced in size for summarization purpose but preserving key vital information and retaining the meaning of the original document. This study presents the Latent Dirichlet Allocation (LDA) approach used to perform topic modelling from summarised medical science journal articles with topics related to genes and diseases. In this study, PyLDAvis web-based interactive visualization tool was used to visualise the selected topics. The visualisation provides an overarching view of the main topics while allowing and attributing deep meaning to the prevalence individual topic. This study presents a novel approach to summarization of single and multiple documents. The results suggest the terms ranked purely by considering their probability of the topic prevalence within the processed document using extractive summarization technique. PyLDAvis visualization describes the flexibility of exploring the terms of the topics' association to the fitted LDA model. The topic modelling result shows prevalence within topics 1 and 2. This association reveals that there is similarity between the terms in topic 1 and 2 in this study. The efficacy of the LDA and the extractive summarization methods were measured using Latent Semantic Analysis (LSA) and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics to evaluate the reliability and validity of the model.
翻译:随着大数据挖掘和海量文本分析在当代的兴起与普及,自动文本摘要在从文档中提取和检索重要信息方面变得日益重要。本研究从单文档和多文档两个视角探讨自动文本摘要问题。摘要是一项将长篇文本压缩为简短摘要版本的任务,在保留原始文档核心关键信息与语义内涵的前提下精简文本篇幅。本研究采用潜在狄利克雷分配(LDA)方法对医学科学期刊论文进行主题建模,主题涵盖基因与疾病相关领域。研究使用基于Web的交互式可视化工具PyLDAvis对选定主题进行可视化展示,该可视化既能提供主要主题的全局概览,又能深入解析各主题的分布特征。本研究提出一种新颖的单文档与多文档摘要方法。结果表明,基于提取式摘要技术,通过单纯考虑词汇在处理文档中主题分布概率进行排序的术语具有显著效果。PyLDAvis可视化展示了探索术语与拟合LDA模型主题关联的灵活性。主题建模结果显示主题1与主题2存在显著分布,这种关联揭示了这两个主题中术语的相似性。研究采用潜在语义分析(LSA)和面向召回率的摘要评估(ROUGE)指标对LDA与提取式摘要方法的效能进行度量,以评估模型的可靠性与有效性。