Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI systems evolve into autonomous agents deployed in open environments and increasingly connected to payments or assets, the operational meaning of trust shifts to end-to-end outcomes: whether an agent completes tasks, follows user intent, and avoids failures that cause material or psychological harm. These risks are fundamentally product-level and cannot be eliminated by technical safeguards alone because agent behavior is inherently stochastic. To address this gap between model-level reliability and user-facing assurance, we propose a complementary framework based on risk management. Drawing inspiration from financial underwriting, we introduce the \textbf{Agentic Risk Standard (ARS)}, a payment settlement standard for AI-mediated transactions. ARS integrates risk assessment, underwriting, and compensation into a single transaction framework that protects users when interacting with agents. Under ARS, users receive predefined and contractually enforceable compensation in cases of execution failure, misalignment, or unintended outcomes. This shifts trust from an implicit expectation about model behavior to an explicit, measurable, and enforceable product guarantee. We also present a simulation study analyzing the social benefits of applying ARS to agentic transactions. ARS's implementation can be found at https://github.com/t54-labs/AgenticRiskStandard.
翻译:可信赖AI的前期研究侧重于模型内部属性,如偏见缓解、对抗鲁棒性和可解释性。随着AI系统演变为部署在开放环境中的自主代理,并日益与支付或资产关联,信任的操作性涵义转向端到端的成果:代理能否完成任务、遵循用户意图并避免造成物质或心理伤害的故障。这些风险本质上是产品级的,无法仅通过技术保障消除,因为代理行为天生具备随机性。为弥合模型级可靠性与面向用户的保障之间的差距,我们提出一种基于风险管理的补充框架。借鉴金融承保思路,我们引入**代理风险标准(ARS)**,一种针对AI中介交易的支付结算标准。ARS将风险评估、承保与补偿整合为单一交易框架,在用户与代理交互时提供保护。在该标准下,用户在出现执行失败、目标偏离或意外结果时,可获得预定义且具备合同约束力的补偿。此举将信任从对模型行为的隐性期望,转变为显性、可度量且可执行的产品承诺。我们还通过模拟研究分析了将ARS应用于代理交易的社会效益。ARS的实现代码可见于https://github.com/t54-labs/AgenticRiskStandard。