Despite the promise of RLHF in aligning LLMs with human preferences, it often leads to superficial alignment, prioritizing stylistic changes over improving downstream performance of LLMs. Underspecified preferences could obscure directions to align the models. Lacking exploration restricts identification of desirable outputs to improve the models. To overcome these challenges, we propose a novel framework: Reinforcement Learning from Reflective Feedback (RLRF), which leverages fine-grained feedback based on detailed criteria to improve the core capabilities of LLMs. RLRF employs a self-reflection mechanism to systematically explore and refine LLM responses, then fine-tuning the models via a RL algorithm along with promising responses. Our experiments across Just-Eval, Factuality, and Mathematical Reasoning demonstrate the efficacy and transformative potential of RLRF beyond superficial surface-level adjustment.
翻译:尽管RLHF(基于人类反馈的强化学习)在使大语言模型与人类偏好对齐方面展现出潜力,但其往往导致表面化对齐——更侧重于风格化调整而非提升模型的下游性能。欠规范偏好可能模糊模型对齐方向,而探索不足则制约了改进模型所需的高质量输出识别。为克服这些挑战,我们提出一种新型框架——基于反思反馈的强化学习(RLRF),该框架利用基于精细化标准的细粒度反馈来提升大语言模型的核心能力。RLRF通过自我反思机制系统性地探索与优化模型响应,随后结合优质响应,采用强化学习算法对模型进行微调。我们在Just-Eval评估、事实性任务与数学推理上的实验表明,RLRF能够超越表层调整,展现出显著的有效性与变革潜力。