Recent studies show that instruction tuning and learning from human feedback improve the abilities of large language models (LMs) dramatically. While these tuning methods can make models generate high-quality text, we conjecture that more implicit cognitive biases may arise in these fine-tuned models. Our work provides evidence that these fine-tuned models exhibit biases that were absent or less pronounced in their pretrained predecessors. We examine the extent of this phenomenon in three cognitive biases - the decoy effect, the certainty effect, and the belief bias - all of which are known to influence human decision-making and reasoning. Our findings highlight the presence of these biases in various models, especially those that have undergone instruction tuning, such as Flan-T5, GPT3.5, and GPT4. This research constitutes a step toward comprehending cognitive biases in instruction-tuned LMs, which is crucial for the development of more reliable and unbiased language models.
翻译:近期研究表明,指令微调与基于人类反馈的学习显著提升了大型语言模型的能力。尽管这类微调方法能使模型生成高质量文本,但我们推测这些经过微调的模型可能隐含更细微的认知偏见。本研究证实,相较于预训练阶段的前代模型,经微调的模型表现出原本不存在或较不显著的偏见。我们针对三种已知影响人类决策与推理的认知偏见——诱饵效应、确定性效应与信念偏见——系统考察了该现象的程度。研究发现表明,不同模型(尤其经过指令微调的Flan-T5、GPT3.5与GPT4)均存在这些偏见。本项研究为理解指令微调语言模型中的认知偏见迈出了重要一步,这对开发更可靠、更无偏的语言模型至关重要。