AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or on self-reported indicators, rather than how task effort is distributed between users and tools. Here, we introduce offloading score, a measure of reliance that quantifies the fraction of cognitive effort offloaded to an AI tool. Offloading Score is simulation-based -- we construct a counterfactual workflow by estimating how the user would have completed the task without the tool, and then computing the fraction of steps saved by using the tool. We validate offloading score through intrinsic evaluations of metric validity, and a controlled user study ($n=40$) with developers performing programming tasks using AI tools. We vary time pressure to test whether reliance measures capture the known increase in reliance under time pressure. We show that offloading score detects significantly higher reliance in time-constrained settings ($+43\%$, $p=0.018$), while usage-based and self-reported baseline measures of reliance do not distinguish the conditions. We complement this with descriptive insights showing that higher reliance manifests as greater delegation of subtasks to the tool and more direct reuse of AI outputs. Finally, we demonstrate an approach of using offloading score in combination with target outcomes of a task (e.g., code understanding) to identify when reliance may be (in)appropriate. Our framework offers two contributions: an instrument users can apply to measure and reflect on their own reliance, and a quantitative signal that agent designers can utilize to mitigate overreliance.
翻译:人工智能工具正日益融入现实工作流程。然而,现有衡量对这些工具依赖程度的指标主要关注AI输出的采用率或自我报告指标,而非任务工作量在用户与工具之间的分配方式。本文提出"离线评分"这一依赖度衡量指标,可量化用户将认知工作量转移给AI工具的比例。该评分基于模拟方法——通过估算用户在没有工具辅助时完成任务的路径构建反事实工作流程,进而计算使用工具所节省的步骤占比。我们通过内在指标有效性评估以及一项由开发人员使用AI工具完成编程任务的受控用户研究(n=40)对离线评分进行验证。通过改变时间压力条件,检验依赖度指标能否捕捉到已知的时间压力下依赖程度升高现象。研究表明,离线评分能显著检测到时间限制情境下依赖度升高(+43%,p=0.018),而基于使用频率和自述报告的基线指标则无法区分不同条件。我们进一步通过描述性分析揭示:更高依赖度表现为将更多子任务授权给工具执行,以及更直接地复用AI输出结果。最后,我们展示了一种结合任务目标结果(如代码理解)使用离线评分的方法,用以识别依赖行为是否适当。本框架可提供两方面的贡献:一是供用户测量与反思自身依赖程度的工具,二是供智能体设计者用于缓解过度依赖现象的量化信号。