In recent years, large language models (LLMs) have shown remarkable capabilities at scale, particularly at generating text conditioned on a prompt. In our work, we investigate the use of LLMs to augment training data of small language models~(SLMs) with automatically generated counterfactual~(CF) instances -- i.e. minimally altered inputs -- in order to improve out-of-domain~(OOD) performance of SLMs in the extractive question answering~(QA) setup. We show that, across various LLM generators, such data augmentation consistently enhances OOD performance and improves model calibration for both confidence-based and rationale-augmented calibrator models. Furthermore, these performance improvements correlate with higher diversity of CF instances in terms of their surface form and semantic content. Finally, we show that CF augmented models which are easier to calibrate also exhibit much lower entropy when assigning importance, indicating that rationale-augmented calibrators prefer concise explanations.
翻译:近年来,大语言模型(LLMs)在规模化条件下展现出显著能力,尤其在基于提示生成文本方面。本研究探索利用LLMs通过自动生成反事实实例(即最小化修改的输入)来增强小语言模型(SLMs)的训练数据,旨在提升抽取式问答场景中SLMs的领域外性能。我们证明:无论采用何种LLM生成器,此类数据增强均能持续提升领域外性能,并优化基于置信度和基于理由增强的校准器模型的校准效果。此外,这些性能改进与反事实实例在表层形式和语义内容上的多样性呈正相关。最后,我们发现,更易校准的反事实增强模型在重要性分配时具有更低的熵值,这表明基于理由增强的校准器更偏好简洁的解释。