Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural language on a range of tasks. As LLMs increasingly influence consequential decisions across diverse domains from healthcare to personal advice, the risk of overreliance -- relying on LLMs beyond their capabilities -- grows. This paper argues that measuring and mitigating overreliance must become central to LLM research and deployment. First, we consolidate risks from overreliance at both the individual and societal levels, including high-stakes errors, governance challenges, and cognitive deskilling. Then, we explore LLM characteristics, system design features, and user cognitive biases that together raise serious and unique concerns about overreliance on LLMs in practice. We also examine historical approaches for measuring overreliance, identifying three important gaps and proposing three promising directions to improve measurement. Finally, we propose mitigation strategies that can be pursued to ensure LLMs augment rather than undermine human capabilities.
翻译:大型语言模型(LLM)通过扮演协作式“思想伙伴”角色,能够在自然语言交互中更流畅地完成多项任务,从而区别于以往的技术。随着LLM在从医疗健康到个人建议等不同领域日益影响关键决策,过度依赖的风险——即超出模型实际能力的依赖行为——也随之增长。本文论证了衡量并缓解过度依赖必须成为LLM研究与部署的核心议题。首先,我们从个体与社会两个层面整合过度依赖引发的风险,包括高风险的决策失误、治理挑战与认知能力退化。随后,我们探讨了LLM特性、系统设计特征及用户认知偏差,这些因素在实践中共同构成了对LLM过度依赖的独特而严峻的隐忧。我们还梳理了过往衡量过度依赖的方法,识别出三个关键缺口,并提出三大可行方向以改进测量。最后,我们提出一系列缓解策略,确保LLM能够增强而非削弱人类能力。