Style analysis, which is relatively a less explored topic, enables several interesting applications. For instance, it allows authors to adjust their writing style to produce a more coherent document in collaboration. Similarly, style analysis can also be used for document provenance and authentication as a primary step. In this paper, we propose an ensemble-based text-processing framework for the classification of single and multi-authored documents, which is one of the key tasks in style analysis. The proposed framework incorporates several state-of-the-art text classification algorithms including classical Machine Learning (ML) algorithms, transformers, and deep learning algorithms both individually and in merit-based late fusion. For the merit-based late fusion, we employed several weight optimization and selection methods to assign merit-based weights to the individual text classification algorithms. We also analyze the impact of the characters on the task that are usually excluded in NLP applications during pre-processing by conducting experiments on both clean and un-clean data. The proposed framework is evaluated on a large-scale benchmark dataset, significantly improving performance over the existing solutions.
翻译:文体分析是一个相对较少被探索的研究领域,却能够支撑多项有趣的应用。例如,它使作者能够调整写作风格,从而在协作中生成更连贯的文档;同样,文体分析也可作为文档溯源与认证的首要步骤。本文提出了一种基于集成学习的文本处理框架,用于对单人作者和多人作者文档进行分类,这是文体分析中的关键任务之一。该框架整合了多种当前最先进的文本分类算法,包括经典机器学习算法、Transformer模型以及深度学习算法,并分别采用独立运行和基于权重优化的后期融合策略。针对基于权重优化的后期融合,我们采用多种权重优化与选择方法,为各文本分类算法分配基于贡献度的权重。此外,我们通过分别在清洁数据与未清洁数据上的实验,分析了通常在自然语言处理预处理阶段被排除的字符对该任务的影响。该框架在大型基准数据集上进行了评估,显著提升了现有解决方案的性能。