While state-of-the-art language models excel at the style transfer task, current work does not address explainability of style transfer systems. Explanations could be generated using large language models such as GPT-3.5 and GPT-4, but the use of such complex systems is inefficient when smaller, widely distributed, and transparent alternatives are available. We propose a framework to augment and improve a formality style transfer dataset with explanations via model distillation from ChatGPT. To further refine the generated explanations, we propose a novel way to incorporate scarce expert human feedback using in-context learning (ICLEF: In-Context Learning from Expert Feedback) by prompting ChatGPT to act as a critic to its own outputs. We use the resulting dataset of 9,960 explainable formality style transfer instances (e-GYAFC) to show that current openly distributed instruction-tuned models (and, in some settings, ChatGPT) perform poorly on the task, and that fine-tuning on our high-quality dataset leads to significant improvements as shown by automatic evaluation. In human evaluation, we show that models much smaller than ChatGPT fine-tuned on our data align better with expert preferences. Finally, we discuss two potential applications of models fine-tuned on the explainable style transfer task: interpretable authorship verification and interpretable adversarial attacks on AI-generated text detectors.
翻译:摘要:尽管最先进的语言模型在风格迁移任务上表现出色,但当前研究并未解决风格迁移系统的可解释性问题。虽然可以利用GPT-3.5和GPT-4等大型语言模型生成解释,但在存在更小、更广泛分布且透明的替代方案时,使用这类复杂系统效率较低。我们提出一种框架,通过从ChatGPT进行模型蒸馏,为形式化风格迁移数据集增强并补充解释。为进一步优化生成的解释,我们提出一种创新方法,利用稀缺的人类专家反馈进行上下文学习(ICLEF:基于专家反馈的上下文学习),通过提示ChatGPT充当自身输出的批评者。我们利用由此生成的9960个可解释形式化风格迁移实例数据集(e-GYAFC)证明,当前公开分布的指令微调模型(在某些设置下包括ChatGPT)在该任务上表现不佳,而基于我们高质量数据集的微调则通过自动评估显示出显著改进。在人工评估中,我们证明基于我们的数据微调后、规模远小于ChatGPT的模型能更好地对齐专家偏好。最后,我们探讨了基于可解释风格迁移任务微调模型的两项潜在应用:可解释的作者身份验证以及针对AI生成文本检测器的可解释对抗攻击。