Fine-tuning language models on tasks with instructions has demonstrated potential in facilitating zero-shot generalization to unseen tasks. In this paper, we introduce a straightforward yet effective method for enhancing instruction tuning by employing symbolic tasks. Compared to crowdsourced human tasks or model-generated tasks, symbolic tasks present a unique advantage as they can be easily generated in vast quantities, theoretically providing an infinite supply of high-quality training instances. To explore the potential of symbolic tasks, we carry out an extensive case study on the representative symbolic task of SQL execution. Empirical results on various benchmarks validate that the integration of SQL execution leads to significant improvements in zero-shot scenarios, particularly in table reasoning. Notably, our 3B model surpasses both the 175B GPT-3 and ChatGPT in zero-shot table reasoning across four benchmarks. Furthermore, experimental results on BBH (27 tasks) and MMLU (57 tasks) reveal that language models can be enhanced through symbolic tasks without compromising their generality. We hope that our paper serves as a catalyst, inspiring increased efforts to incorporate symbolic tasks in instruction tuning.
翻译:对带有指令的任务进行语言模型微调已展现出促进对未见任务零样本泛化的潜力。本文提出一种简单而有效的方法,即通过使用符号任务增强指令微调。与人工众包任务或模型生成任务相比,符号任务具有独特优势:它们可以轻松大规模生成,理论上提供无限量的高质量训练实例。为探索符号任务的潜力,我们以代表性符号任务——SQL执行为案例开展广泛研究。在多个基准上的实证结果验证了整合SQL执行可显著提升零样本场景下的性能,尤其在表格推理任务中。值得注意的是,我们的3B模型在四个基准的零样本表格推理上超越了175B的GPT-3和ChatGPT。此外,BBH(27个任务)和MMLU(57个任务)上的实验结果表明,语言模型可通过符号任务增强而不损害其通用性。我们希望本文能成为催化剂,激励更多研究将符号任务纳入指令微调中。