Advancements in large language models (LLMs) have demonstrated remarkable capabilities across a diverse range of applications. These models excel in generating text completions that are contextually coherent and cover an extensive array of subjects. However, the vast datasets required for their training make aligning response styles during the pretraining and instruction tuning phases challenging. Consequently, an additional alignment phase is typically employed, wherein the model is further trained with human preference data to better align its outputs with human expectations. While this process doesn't introduce new capabilities per se, it does accentuate generation styles innate to the model. This paper explores the utilization of counterfactual prompting within the framework of Direct Preference Optimization (DPO) to align the model's style without relying on human intervention. We demonstrate that this method effectively instils desirable behaviour, mitigates undesirable ones, and encourages the model to disregard inappropriate instructions. Our findings suggest that counterfactual prompting with DPO presents a low-resource way to fine-tune LLMs to meet the demands for responsible and ethically aligned AI systems.
翻译:大语言模型(LLMs)的进步已在众多应用领域展现出卓越能力,这些模型能生成上下文连贯且覆盖广泛主题的文本。然而,训练所需的海量数据集使得在预训练与指令微调阶段统一响应风格颇具挑战。为此,通常需要额外的对齐阶段:利用人类偏好数据对模型进行进一步训练,使其输出更符合人类预期。该过程虽未引入新能力,但能强化模型固有的生成风格。本文探索在直接偏好优化框架中运用反事实提示技术,实现无需人类干预的模型风格对齐。实验表明,该方法能有效培养期望行为、抑制不良倾向,并促使模型拒绝不当指令。研究结果揭示,基于反事实提示的DPO技术为大语言模型的低资源微调提供了可行路径,有助于构建负责任且符合伦理的AI系统。