Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI systems evolve into autonomous agents deployed in open environments and increasingly connected to payments or assets, the operational meaning of trust shifts to end-to-end outcomes: whether an agent completes tasks, follows user intent, and avoids failures that cause material or psychological harm. These risks are fundamentally product-level and cannot be eliminated by technical safeguards alone because agent behavior is inherently stochastic. To address this gap between model-level reliability and user-facing assurance, we propose a complementary framework based on risk management. Drawing inspiration from financial underwriting, we introduce the \textbf{Agentic Risk Standard (ARS)}, a payment settlement standard for AI-mediated transactions. ARS integrates risk assessment, underwriting, and compensation into a single transaction framework that protects users when interacting with agents. Under ARS, users receive predefined and contractually enforceable compensation in cases of execution failure, misalignment, or unintended outcomes. This shifts trust from an implicit expectation about model behavior to an explicit, measurable, and enforceable product guarantee. We also present a simulation study analyzing the social benefits of applying ARS to agentic transactions. ARS's implementation can be found at https://github.com/t54-labs/AgenticRiskStandard.


翻译:可信赖AI的前期研究侧重于模型内部属性,如偏见缓解、对抗鲁棒性和可解释性。随着AI系统演变为部署在开放环境中的自主代理,并日益与支付或资产关联,信任的操作性涵义转向端到端的成果:代理能否完成任务、遵循用户意图并避免造成物质或心理伤害的故障。这些风险本质上是产品级的,无法仅通过技术保障消除,因为代理行为天生具备随机性。为弥合模型级可靠性与面向用户的保障之间的差距,我们提出一种基于风险管理的补充框架。借鉴金融承保思路,我们引入**代理风险标准(ARS)**,一种针对AI中介交易的支付结算标准。ARS将风险评估、承保与补偿整合为单一交易框架,在用户与代理交互时提供保护。在该标准下,用户在出现执行失败、目标偏离或意外结果时,可获得预定义且具备合同约束力的补偿。此举将信任从对模型行为的隐性期望,转变为显性、可度量且可执行的产品承诺。我们还通过模拟研究分析了将ARS应用于代理交易的社会效益。ARS的实现代码可见于https://github.com/t54-labs/AgenticRiskStandard。

0
下载
关闭预览

相关内容

《军事应用中的AI:建立信任》最新报告
专知会员服务
25+阅读 · 2025年12月29日
智能金融稳步前行:构建负责任的可信大模型
专知会员服务
22+阅读 · 2024年10月8日
【资源推荐】AI可解释性资源汇总
专知
47+阅读 · 2019年4月24日
【智能金融】机器学习在反欺诈中应用
产业智能官
35+阅读 · 2019年3月15日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
综述 | 终端智能体:命令行环境中的 AI Agents
专知会员服务
3+阅读 · 8月24日
综述 | 多模态智能体框架基础与前沿
专知会员服务
3+阅读 · 8月24日
《自适应无人机集群网络开发》
专知会员服务
6+阅读 · 8月24日
无人机与作战飞机最佳精确目标定位技术
专知会员服务
5+阅读 · 8月24日
博士论文 | 大动作空间中的在线与离线策略学习
论文 | 全球负责任 AI 指数 2026 方法论
专知会员服务
7+阅读 · 8月23日
超致命战场空间中的战术通信生存能力
专知会员服务
5+阅读 · 8月23日
相关VIP内容
《军事应用中的AI:建立信任》最新报告
专知会员服务
25+阅读 · 2025年12月29日
智能金融稳步前行:构建负责任的可信大模型
专知会员服务
22+阅读 · 2024年10月8日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员