Text simplification aims to make the text easier to understand by applying rewriting transformations. There has been very little research on Chinese text simplification for a long time. The lack of generic evaluation data is an essential reason for this phenomenon. In this paper, we introduce MCTS, a multi-reference Chinese text simplification dataset. We describe the annotation process of the dataset and provide a detailed analysis of it. Furthermore, we evaluate the performance of some unsupervised methods and advanced large language models. We hope to build a basic understanding of Chinese text simplification through the foundational work and provide references for future research. We release our data at https://github.com/blcuicall/mcts.
翻译:文本简化旨在通过重写转换使文本更易理解。长期以来,针对中文文本简化的研究非常有限。缺乏通用的评估数据是导致这一现象的重要原因。本文介绍了MCTS——一个多参考中文文本简化数据集。我们描述了该数据集的标注流程并进行了详细分析。此外,我们评估了一些无监督方法和先进大语言模型的性能。希望通过这项基础工作,建立对中文文本简化的基本理解,并为未来研究提供参考。我们已在https://github.com/blcuicall/mcts公开该数据集。