Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL training environment, which reduces deployment overhead in Docker-level environments by two orders of magnitude. Finally, we deploy AgentDoG 1.5 as a training-free online guardrail for real-time safety moderation. Extensive experimental results indicate that AgentDoG 1.5 achieves state-of-the-art performance in diverse and complex interactive agentic scenarios. All models and datasets are openly released.


翻译:现代开放世界智能体(如OpenClaw)展现出强大的跨环境执行能力,但也带来了全新的安全风险源。同时,前沿AI模型大幅降低了攻击门槛,导致现有智能体对齐框架难以满足真实部署需求。为应对这些新兴威胁,我们提出了一种轻量级且可扩展的智能体安全对齐框架。具体而言,我们更新了智能体安全分类体系以涵盖来自Codex和OpenClaw执行场景的涌现风险,并进一步构建了基于影响力函数净化的分类引导数据引擎,仅使用约1000个样本即可训练轻量级AgentDoG 1.5变体(参数量0.8B/2B/4B/8B),达到与领先闭源模型(如GPT-5.4)相近的性能。基于AgentDoG 1.5,我们搭建了高效的智能体安全SFT和RL训练环境,将Docker级环境的部署开销降低两个数量级。最终我们将AgentDoG 1.5部署为免训练的在线防护栏实现实时安全审查。大量实验表明,AgentDoG 1.5在多样化的复杂交互式智能体场景中达到了最先进性能。所有模型与数据集均已开源发布。

0
下载
关闭预览

相关内容

AgentOps综述:智能体系统运维框架
专知会员服务
24+阅读 · 6月4日
通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
AI智能体时代大模型安全风险与攻防新挑战
专知会员服务
15+阅读 · 2月27日
AI 智能体系统:体系架构、应用场景及评估范式
AI智能体基础设施
专知会员服务
44+阅读 · 2025年7月12日
人工智能如何增强军事监控与边境安全
专知会员服务
22+阅读 · 2025年3月20日
机密计算保障人工智能系统安全研究报告
专知会员服务
20+阅读 · 2025年1月20日
专知会员服务
64+阅读 · 2021年7月5日
《人工智能安全框架(2020年)》白皮书,68页pdf
专知会员服务
167+阅读 · 2021年1月9日
面向多智能体博弈对抗的对手建模框架
专知
18+阅读 · 2022年9月28日
重磅!AI框架发展白皮书(2022年),44页pdf
专知
28+阅读 · 2022年2月27日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
浅谈群体智能——新一代AI的重要方向
中国科学院自动化研究所
44+阅读 · 2019年10月16日
基于车路协同的群体智能协同
智能交通技术
10+阅读 · 2019年1月23日
智能时代如何构建金融反欺诈体系?
数据猿
12+阅读 · 2018年3月26日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
AgentOps综述:智能体系统运维框架
专知会员服务
24+阅读 · 6月4日
通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
AI智能体时代大模型安全风险与攻防新挑战
专知会员服务
15+阅读 · 2月27日
AI 智能体系统:体系架构、应用场景及评估范式
AI智能体基础设施
专知会员服务
44+阅读 · 2025年7月12日
人工智能如何增强军事监控与边境安全
专知会员服务
22+阅读 · 2025年3月20日
机密计算保障人工智能系统安全研究报告
专知会员服务
20+阅读 · 2025年1月20日
专知会员服务
64+阅读 · 2021年7月5日
《人工智能安全框架(2020年)》白皮书,68页pdf
专知会员服务
167+阅读 · 2021年1月9日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员