Complex text is a major barrier for many citizens when accessing public information and knowledge. While often done manually, Text Simplification is a key Natural Language Processing task that aims for reducing the linguistic complexity of a text while preserving the original meaning. Recent advances in Generative Artificial Intelligence (AI) have enabled automatic text simplification both on the lexical and syntactical levels. However, as applications often focus on English, little is understood about the effectiveness of Generative AI techniques on low-resource languages such as Dutch. For this reason, we carry out empirical studies to understand the benefits and limitations of applying generative technologies for text simplification and provide the following outcomes: 1) the design and implementation for a configurable text simplification pipeline that orchestrates state-of-the-art generative text simplification models, domain and reader adaptation, and visualisation modules; 2) insights and lessons learned, showing the strengths of automatic text simplification while exposing the challenges in handling cultural and commonsense knowledge. These outcomes represent a first step in the exploration of Dutch text simplification and shed light on future endeavours both for research and practice.
翻译:复杂文本是许多公民在获取公共信息和知识时面临的主要障碍。文本简化作为一项关键的自然语言处理任务,旨在降低文本的语言复杂性,同时保留原始含义,该任务目前通常由人工完成。生成式人工智能的最新进展已在词汇和句法层面实现了自动文本简化。然而,由于相关应用通常聚焦于英语,人们对生成式人工智能技术在处理荷兰语等低资源语言时的有效性知之甚少。为此,我们开展实证研究,以理解生成技术应用于文本简化的优势与局限,并取得以下成果:1)设计并实现了一个可配置的文本简化流水线,该流水线整合了最先进的生成式文本简化模型、领域与读者自适应模块以及可视化模块;2)获取的洞见与经验教训表明,自动文本简化在展现优势的同时,仍面临处理文化知识与常识性知识的挑战。这些成果代表了荷兰语文本简化探索的初步尝试,并为未来的研究与实践指明了方向。