In recent years, with the rapid development of information on the Internet, the number of complex texts and documents has increased exponentially, which requires a deeper understanding of deep learning methods in order to accurately classify texts using deep learning techniques, and thus deep learning methods have become increasingly important in text classification. Text classification is a class of tasks that automatically classifies a set of documents into multiple predefined categories based on their content and subject matter. Thus, the main goal of text classification is to enable users to extract information from textual resources and process processes such as retrieval, classification, and machine learning techniques together in order to classify different categories. Many new techniques of deep learning have already achieved excellent results in natural language processing. The success of these learning algorithms relies on their ability to understand complex models and non-linear relationships in data. However, finding the right structure, architecture, and techniques for text classification is a challenge for researchers. This paper introduces deep learning-based text classification algorithms, including important steps required for text classification tasks such as feature extraction, feature reduction, and evaluation strategies and methods. At the end of the article, different deep learning text classification methods are compared and summarized.
翻译:近年来,随着互联网信息的快速发展,复杂文本和文档的数量呈指数级增长,这要求对深度学习方法有更深入的理解,以便利用深度学习技术准确地进行文本分类,因此深度学习方法在文本分类中日益重要。文本分类是一类根据文档的内容和主题自动将其归入多个预定义类别的任务。因此,文本分类的主要目标是使用户能够从文本资源中提取信息,并将检索、分类与机器学习技术等处理过程相结合,以区分不同的类别。许多新的深度学习技术已经在自然语言处理中取得了优异成果。这些学习算法的成功依赖于它们理解数据中复杂模型和非线性关系的能力。然而,为文本分类找到合适的结构、架构和技术对研究者而言仍是一项挑战。本文介绍了基于深度学习的文本分类算法,包括文本分类任务所需的重要步骤,如特征提取、特征降维以及评估策略与方法。文章最后对不同深度学习文本分类方法进行了比较与总结。