Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL training environment, which reduces deployment overhead in Docker-level environments by two orders of magnitude. Finally, we deploy AgentDoG 1.5 as a training-free online guardrail for real-time safety moderation. Extensive experimental results indicate that AgentDoG 1.5 achieves state-of-the-art performance in diverse and complex interactive agentic scenarios. All models and datasets are openly released.


翻译:现代开放世界智能体(如OpenClaw)展现出强大的跨环境执行能力,但也带来了全新的安全风险源。同时,前沿AI模型大幅降低了攻击门槛,导致现有智能体对齐框架难以满足真实部署需求。为应对这些新兴威胁,我们提出了一种轻量级且可扩展的智能体安全对齐框架。具体而言,我们更新了智能体安全分类体系以涵盖来自Codex和OpenClaw执行场景的涌现风险,并进一步构建了基于影响力函数净化的分类引导数据引擎,仅使用约1000个样本即可训练轻量级AgentDoG 1.5变体(参数量0.8B/2B/4B/8B),达到与领先闭源模型(如GPT-5.4)相近的性能。基于AgentDoG 1.5,我们搭建了高效的智能体安全SFT和RL训练环境,将Docker级环境的部署开销降低两个数量级。最终我们将AgentDoG 1.5部署为免训练的在线防护栏实现实时安全审查。大量实验表明,AgentDoG 1.5在多样化的复杂交互式智能体场景中达到了最先进性能。所有模型与数据集均已开源发布。

0
下载
关闭预览

相关内容

AgentOps综述:智能体系统运维框架
专知会员服务
25+阅读 · 6月4日
通用智能体评估的逻辑架构
专知会员服务
23+阅读 · 2月28日
AI智能体时代大模型安全风险与攻防新挑战
专知会员服务
16+阅读 · 2月27日
AI 智能体系统:体系架构、应用场景及评估范式
AI智能体基础设施
专知会员服务
44+阅读 · 2025年7月12日
人工智能如何增强军事监控与边境安全
专知会员服务
22+阅读 · 2025年3月20日
机密计算保障人工智能系统安全研究报告
专知会员服务
20+阅读 · 2025年1月20日
专知会员服务
65+阅读 · 2021年7月5日
《人工智能安全框架(2020年)》白皮书,68页pdf
专知会员服务
167+阅读 · 2021年1月9日
面向多智能体博弈对抗的对手建模框架
专知
19+阅读 · 2022年9月28日
重磅!AI框架发展白皮书(2022年),44页pdf
专知
28+阅读 · 2022年2月27日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
浅谈群体智能——新一代AI的重要方向
中国科学院自动化研究所
44+阅读 · 2019年10月16日
基于车路协同的群体智能协同
智能交通技术
10+阅读 · 2019年1月23日
智能时代如何构建金融反欺诈体系?
数据猿
12+阅读 · 2018年3月26日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
专知会员服务
1+阅读 · 今天15:40
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关VIP内容
AgentOps综述:智能体系统运维框架
专知会员服务
25+阅读 · 6月4日
通用智能体评估的逻辑架构
专知会员服务
23+阅读 · 2月28日
AI智能体时代大模型安全风险与攻防新挑战
专知会员服务
16+阅读 · 2月27日
AI 智能体系统:体系架构、应用场景及评估范式
AI智能体基础设施
专知会员服务
44+阅读 · 2025年7月12日
人工智能如何增强军事监控与边境安全
专知会员服务
22+阅读 · 2025年3月20日
机密计算保障人工智能系统安全研究报告
专知会员服务
20+阅读 · 2025年1月20日
专知会员服务
65+阅读 · 2021年7月5日
《人工智能安全框架(2020年)》白皮书,68页pdf
专知会员服务
167+阅读 · 2021年1月9日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员