Recent advancements in Large Language Models (LLMs) have expanded the horizons of natural language understanding and generation. Notably, the output control and alignment with the input of LLMs can be refined through instruction tuning. However, as highlighted in several studies, low-quality data in the training set are usually detrimental to instruction tuning, resulting in inconsistent or even misleading LLM outputs. We propose a novel method, termed "reflection-tuning," which addresses the problem by self-improvement and judging capabilities of LLMs. This approach utilizes an oracle LLM to recycle the original training data by introspecting and enhancing the quality of instructions and responses in the data. Extensive experiments on widely used evaluation benchmarks show that LLMs trained with our recycled data outperform those trained with existing datasets in various benchmarks.
翻译:近年来,大语言模型的进步拓展了自然语言理解与生成的边界。值得注意的是,通过指令微调可以优化大语言模型的输出控制及其与输入的匹配性。然而,多项研究表明,训练集中的低质量数据通常会对指令微调产生不利影响,导致大语言模型输出不一致甚至产生误导。我们提出一种名为"反思微调"的新方法,该方法通过利用大语言模型的自我改进和判断能力来解决这一问题。该方法借助一个"神谕"大语言模型,通过内省并提升数据中指令与响应的质量,对原始训练数据进行回收利用。在广泛使用的评估基准上进行的多项实验表明,使用我们回收数据训练的大语言模型,在各类基准测试中的表现均优于使用现有数据集训练的模型。