An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause harm and must be taken offline, curtailing any future interaction. Imitating old behavior is safe, but excessive conservatism discourages exploration. How much behavior change is too much? We show how to use any safe reference policy as a probabilistic regulator for any optimized but untested policy. Conformal calibration on data from the safe policy determines how aggressively the new policy can act, while provably enforcing the user's declared risk tolerance. Unlike conservative optimization methods, we do not assume the user has identified the correct model class nor tuned any hyperparameters. Unlike previous conformal methods, our theory provides finite-sample guarantees even for non-monotonic bounded loss functions. Our experiments on applications ranging from natural language question answering to biomolecular engineering show that safe exploration is not only possible from the first moment of deployment, but can also improve performance.


翻译:代理必须尝试新行为以进行探索和改进。在高风险环境中,违反安全约束的代理可能造成伤害,且必须被下线处理,从而终止任何未来交互。模仿旧行为是安全的,但过度保守会抑制探索。行为改变多少才算过度?我们展示了如何利用任何安全参考策略作为概率调节器,以约束任何已优化但未经测试的策略。基于安全策略数据的保形校准决定了新策略可以采取行动的激进程度,同时可证明地强制执行用户声明的风险容忍度。与保守优化方法不同,我们既不假设用户已确定正确的模型类别,也不假设其已调整任何超参数。与以往的保形方法不同,我们的理论即使对于非单调有界损失函数也能提供有限样本保证。我们在从自然语言问答到生物分子工程等应用上的实验表明,安全探索不仅可以在部署初始阶段实现,还能提升性能。

0
下载
关闭预览

相关内容

认知优势:人工智能在国家安全决策中的核心作用
专知会员服务
16+阅读 · 2025年8月16日
《视觉Transformers自监督学习机制综述》
专知会员服务
29+阅读 · 2024年9月2日
语言模型在军事和外交决策中的升级风险
专知会员服务
28+阅读 · 2024年6月24日
《大型语言模型保护措施》综述
专知会员服务
29+阅读 · 2024年6月6日
结构保持图transformer综述
专知会员服务
42+阅读 · 2024年2月19日
专知会员服务
10+阅读 · 2020年11月12日
深度学习的下一步:Transformer和注意力机制
云头条
56+阅读 · 2019年9月14日
深入理解BERT Transformer ,不仅仅是注意力机制
大数据文摘
22+阅读 · 2019年3月19日
【智能制造】智能制造的核心——智能决策
产业智能官
12+阅读 · 2018年4月11日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 4月28日
Arxiv
0+阅读 · 4月7日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
7+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
5+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
相关VIP内容
认知优势:人工智能在国家安全决策中的核心作用
专知会员服务
16+阅读 · 2025年8月16日
《视觉Transformers自监督学习机制综述》
专知会员服务
29+阅读 · 2024年9月2日
语言模型在军事和外交决策中的升级风险
专知会员服务
28+阅读 · 2024年6月24日
《大型语言模型保护措施》综述
专知会员服务
29+阅读 · 2024年6月6日
结构保持图transformer综述
专知会员服务
42+阅读 · 2024年2月19日
专知会员服务
10+阅读 · 2020年11月12日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员