Large Language Models (LLMs) have demonstrated remarkable capabilities in various NLP tasks. However, previous works have shown these models are sensitive towards prompt wording, and few-shot demonstrations and their order, posing challenges to fair assessment of these models. As these models become more powerful, it becomes imperative to understand and address these limitations. In this paper, we focus on LLMs robustness on the task of multiple-choice questions -- commonly adopted task to study reasoning and fact-retrieving capability of LLMs. Investigating the sensitivity of LLMs towards the order of options in multiple-choice questions, we demonstrate a considerable performance gap of approximately 13% to 75% in LLMs on different benchmarks, when answer options are reordered, even when using demonstrations in a few-shot setting. Through a detailed analysis, we conjecture that this sensitivity arises when LLMs are uncertain about the prediction between the top-2/3 choices, and specific options placements may favor certain prediction between those top choices depending on the question caused by positional bias. We also identify patterns in top-2 choices that amplify or mitigate the model's bias toward option placement. We found that for amplifying bias, the optimal strategy involves positioning the top two choices as the first and last options. Conversely, to mitigate bias, we recommend placing these choices among the adjacent options. To validate our conjecture, we conduct various experiments and adopt two approaches to calibrate LLMs' predictions, leading to up to 8 percentage points improvement across different models and benchmarks.
翻译:大语言模型(LLMs)在各种自然语言处理任务中展现出卓越的能力。然而,先前研究表明,这些模型对提示措辞、少样本示例及其顺序十分敏感,这给模型的公平评估带来了挑战。随着这些模型日益强大,理解和解决这些局限性变得至关重要。本文聚焦于LLMs在多项选择题任务中的鲁棒性——这是研究LLMs推理与事实检索能力的常用任务。通过探究LLMs对多项选择题选项顺序的敏感性,我们发现在不同基准测试中,即使采用少样本设置下的示例,重新排列选项顺序仍会导致模型性能出现约13%到75%的显著差距。通过详细分析,我们推测这种敏感性源于LLMs对前2/3个预测选项的不确定性,且特定选项排列方式可能因位置偏差而对前几个选项的预测产生偏向性。我们还识别出放大或缓解模型对选项排列位置偏差的前两个选项模式。研究发现,为放大偏差,最优策略是将前两个选项置于首位和末位;反之,为缓解偏差,建议将这些选项置于相邻位置。为验证猜想,我们开展多项实验并采用两种校准LLMs预测的方法,在不同模型和基准测试中实现了高达8个百分点的性能提升。