A growing trend involves integrating human knowledge into learning frameworks, leveraging subtle human feedback to refine AI models. Despite these advances, no comprehensive theoretical framework describing the specific conditions under which human comparisons improve the traditional supervised fine-tuning process has been developed. To bridge this gap, this paper studies the effective use of human comparisons to address limitations arising from noisy data and high-dimensional models. We propose a two-stage "Supervised Fine Tuning+Human Comparison" (SFT+HC) framework connecting machine learning with human feedback through a probabilistic bisection approach. The two-stage framework first learns low-dimensional representations from noisy-labeled data via an SFT procedure, and then uses human comparisons to improve the model alignment. To examine the efficacy of the alignment phase, we introduce a novel concept termed the "label-noise-to-comparison-accuracy" (LNCA) ratio. This paper theoretically identifies the conditions under which the "SFT+HC" framework outperforms pure SFT approach, leveraging this ratio to highlight the advantage of incorporating human evaluators in reducing sample complexity. We validate that the proposed conditions for the LNCA ratio are met in a case study conducted via an Amazon Mechanical Turk experiment.
翻译:将人类知识融入学习框架、利用细微的人类反馈优化AI模型已成为一种日益增长的趋势。尽管取得了这些进展,但尚未建立全面的理论框架来描述人类比较在何种具体条件下能够改进传统的监督微调过程。为填补这一空白,本文研究了有效利用人类比较来应对噪声数据和高维模型带来的局限性问题。我们提出了一种两阶段“监督微调+人类比较”(SFT+HC)框架,通过概率二分法将机器学习与人类反馈相结合。该两阶段框架首先通过SFT过程从带噪声标签的数据中学习低维表示,随后利用人类比较来改进模型对齐。为检验对齐阶段的有效性,我们引入了一个名为“标签噪声-比较精度”(LNCA)比率的新概念。本文从理论上确定了“SFT+HC”框架优于纯SFT方法的条件,利用该比率突出显示了引入人类评估者在降低样本复杂度方面的优势。我们通过亚马逊土耳其机器人实验的案例研究验证了所提出的LNCA比率条件得以满足。