Large-scale language models (LLMs) has shown remarkable capability in various of Natural Language Processing (NLP) tasks and attracted lots of attention recently. However, some studies indicated that large language models fail to achieve promising result beyond the state-of-the-art models in English grammatical error correction (GEC) tasks. In this report, we aim to explore the how large language models perform on Chinese grammatical error correction tasks and provide guidance for future work. We conduct experiments with 3 different LLMs of different model scale on 4 Chinese GEC dataset. Our experimental results indicate that the performances of LLMs on automatic evaluation metrics falls short of the previous sota models because of the problem of over-correction. Furthermore, we also discover notable variations in the performance of LLMs when evaluated on different data distributions. Our findings demonstrates that further investigation is required for the application of LLMs on Chinese GEC task.
翻译:大规模语言模型(LLMs)近年来在多种自然语言处理(NLP)任务中展现出卓越能力,并引起了广泛关注。然而,部分研究表明,在英语语法纠错(GEC)任务中,大规模语言模型未能取得超越当前最优模型(state-of-the-art models)的显著成果。本报告旨在探究大规模语言模型在中文语法纠错任务中的表现,为后续研究提供指导。我们在4个中文GEC数据集上,针对3种不同规模的大语言模型进行了实验。实验结果表明,受过度纠正(over-correction)问题的影响,LLMs在自动评估指标上的表现不及此前的最优模型。此外,我们还发现LLMs在不同数据分布下的评估表现存在显著差异。我们的研究结果说明,LLMs在中文GEC任务中的应用仍需进一步探索。