GPT-3 models are very powerful, achieving high performance on a variety of natural language processing tasks. However, there is a relative lack of detailed published analysis on how well they perform on the task of grammatical error correction (GEC). To address this, we perform experiments testing the capabilities of a GPT-3 model (text-davinci-003) against major GEC benchmarks, comparing the performance of several different prompts, including a comparison of zero-shot and few-shot settings. We analyze intriguing or problematic outputs encountered with different prompt formats. We report the performance of our best prompt on the BEA-2019 and JFLEG datasets using a combination of automatic metrics and human evaluations, revealing interesting differences between the preferences of human raters and the reference-based automatic metrics.
翻译:GPT-3模型功能强大,在多种自然语言处理任务中表现出色。然而,关于其在语法错误纠正(GEC)任务中的性能,目前缺乏详细的公开分析。为此,我们通过实验测试了GPT-3模型(text-davinci-003)在主要GEC基准上的能力,比较了多种不同提示的性能,包括零样本与少样本设置的对比。我们分析了不同提示格式中遇到的具有研究价值或有问题的输出结果。通过结合自动评估指标与人工评价,报告了最佳提示在BEA-2019和JFLEG数据集上的表现,揭示了人工评分员偏好与基于参考的自动指标之间存在的显著差异。