Traditionally, decision support studies how humans use machine learning models to make better decisions. In modern agentic systems, this division of roles is increasingly reversed: AI agents act on behalf of users, while humans and tools becomes support mechanisms around them. This role reversal brings reliability concerns to the forefront, since agentic errors can be consequential and agent behavior must remain aligned with human goals and constraints. Departing from the classical view of decision support, we revisit its two basic principles, the cost--value tradeoff of seeking support and the role of uncertainty quantification, in a setting where AI agents are the central actors. We propose a framework for strategic decision support for AI agents through an optimization problem that minimizes support usage subject to controlling a counterfactual missed-support error: the probability that the agent acts alone on instances where support would have materially improved its output. At the population level, we show that the optimal policy is a threshold rule on the value of support. Building on this structure, we develop an online algorithm that adaptively thresholds such a score and uses randomized exploration to control missed-support error without distributional assumptions. We further introduce a calibration-on-the-fly method that reduces unnecessary support calls online. We instantiate this framework across diverse scenarios, including information gathering, human--AI collaboration, and tool use, showing how each can be modeled through the same strategic decision-support lens. Experiments across these settings show that our method reliably controls the target error while substantially reducing support usage in practice.
翻译:传统上,决策支持研究人类如何利用机器学习模型做出更优决策。在现代智能体系统中,这种角色分工正日益逆转:AI智能体代表用户采取行动,而人类和工具则成为围绕其运转的支持机制。这种角色逆转将可靠性问题推至前沿,因为智能体错误可能造成严重后果,且其行为必须始终与人类目标和约束保持一致。本文脱离决策支持的经典视角,在AI智能体作为核心行动者的情境下,重新审视其两大基本原则——寻求支持的成本-价值权衡以及不确定性量化的作用。我们提出一个面向AI智能体的战略决策支持框架,通过求解优化问题来最小化支持使用频次,同时控制反事实遗漏支持误差:即智能体在支持本可显著改善其输出的实例上单独行动的概率。在总体层面,我们证明最优策略是基于支持价值的阈值规则。基于这一结构,我们开发了一种在线算法,该算法自适应地设定此类得分的阈值,并通过随机探索在无分布假设条件下控制遗漏支持误差。我们进一步引入即时校准方法,以在线方式减少不必要的支持调用。我们将该框架应用于信息收集、人机协作和工具使用等多种场景,展示每个场景如何通过相同的战略决策支持视角进行建模。跨场景实验表明,我们的方法在有效控制目标误差的同时,能在实践中大幅降低支持使用频次。