The volume of information is increasing at an incredible rate with the rapid development of the Internet and electronic information services. Due to time constraints, we don't have the opportunity to read all this information. Even the task of analyzing textual data related to one field requires a lot of work. The text summarization task helps to solve these problems. This article presents an experiment on summarization task for Uzbek language, the methodology was based on text abstracting based on TF-IDF algorithm. Using this density function, semantically important parts of the text are extracted. We summarize the given text by applying the n-gram method to important parts of the whole text. The authors used a specially handcrafted corpus called "School corpus" to evaluate the performance of the proposed method. The results show that the proposed approach is effective in extracting summaries from Uzbek language text and can potentially be used in various applications such as information retrieval and natural language processing. Overall, this research contributes to the growing body of work on text summarization in under-resourced languages.
翻译:随着互联网和电子信息服务的高速发展,信息量正以惊人的速度增长。受时间限制,我们无法阅读所有信息,即便是分析单一领域的文本数据也需要大量工作。文本摘要任务有助于解决这些问题。本文介绍了针对乌兹别克语的摘要任务实验,其方法基于TF-IDF算法的文本摘要抽取。利用该密度函数,可提取文本中语义重要的部分。我们通过对整篇文本中的重要部分应用n-gram方法来实现文本摘要。作者使用专门构建的"学校语料库"(School corpus)评估所提方法的性能。实验结果表明,该方法能有效提取乌兹别克语文本摘要,并有望应用于信息检索和自然语言处理等多个领域。总体而言,本研究丰富了低资源语言文本摘要领域的相关研究成果。