Pretrained large language models (LLMs) are able to solve a wide variety of tasks through transfer learning. Various explainability methods have been developed to investigate their decision making process. TracIn (Pruthi et al., 2020) is one such gradient-based method which explains model inferences based on the influence of training examples. In this paper, we explore the use of TracIn to improve model performance in the parameter-efficient tuning (PET) setting. We develop conversational safety classifiers via the prompt-tuning PET method and show how the unique characteristics of the PET regime enable TracIn to identify the cause for certain misclassifications by LLMs. We develop a new methodology for using gradient-based explainability techniques to improve model performance, G-BAIR: gradient-based automated iterative recovery. We show that G-BAIR can recover LLM performance on benchmarks after manually corrupting training labels. This suggests that influence methods like TracIn can be used to automatically perform data cleaning, and introduces the potential for interactive debugging and relabeling for PET-based transfer learning methods.
翻译:预训练大型语言模型(LLMs)能够通过迁移学习解决多种任务。研究者已开发出多种可解释性方法来探究其决策过程。TracIn(Pruthi等人,2020)是一种基于梯度的解释方法,能够根据训练样本的影响程度解释模型推理。本文探索了在参数高效微调(PET)场景下利用TracIn提升模型性能的方法。我们通过提示调优的PET方法构建了对话安全分类器,并揭示了PET模式的独特性如何使TracIn能够识别大型语言模型某些错误分类的成因。我们提出了一种利用基于梯度的可解释性技术提升模型性能的新方法论——G-BAIR:基于梯度的自动迭代恢复。实验表明,在人为破坏训练标签后,G-BAIR能够恢复模型在基准测试上的性能。这表明TracIn等影响度方法可自动用于数据清洗,并开创了基于PET的迁移学习方法进行交互式调试与标签修正的可能性。