Text summarization is an essential task in natural language processing, and researchers have developed various approaches over the years, ranging from rule-based systems to neural networks. However, there is no single model or approach that performs well on every type of text. We propose a system that recommends the most suitable summarization model for a given text. The proposed system employs a fully connected neural network that analyzes the input content and predicts which summarizer should score the best in terms of ROUGE score for a given input. The meta-model selects among four different summarization models, developed for the Slovene language, using different properties of the input, in particular its Doc2Vec document representation. The four Slovene summarization models deal with different challenges associated with text summarization in a less-resourced language. We evaluate the proposed SloMetaSum model performance automatically and parts of it manually. The results show that the system successfully automates the step of manually selecting the best model.
翻译:文本摘要是自然语言处理中的关键任务,研究人员多年来开发了从规则系统到神经网络的多种方法。然而,目前尚无单一模型或方法能在所有文本类型上表现优异。我们提出一个系统,可根据给定文本推荐最合适的摘要模型。该系统采用全连接神经网络分析输入内容,并预测哪个摘要器能在给定输入的ROUGE评分上表现最佳。该元模型基于输入的不同属性(特别是其Doc2Vec文档表示),在四种为斯洛文尼亚语开发的摘要模型中进行选择。这四种斯洛文尼亚语摘要模型分别应对低资源语言文本摘要中的不同挑战。我们通过自动化方式评估所提出的SloMetaSum模型性能,并对其部分内容进行人工评估。结果表明,该系统成功实现了人工选择最佳模型的自动化步骤。