World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate that world models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain that can result in the deployment of unsafe or otherwise compromised robotic policies despite training on seemingly safe ground truth training data. In contrast to traditional data poisoning techniques which directly implant dangerous trajectories into sold or uploaded datasets, our novel attack methods inject malicious prompts or compromising transition dynamics into visibly safe teleoperated datasets which are only activated once fed through a world model as input. This can result in the generation of synthetic, dangerous robot training trajectories and subsequently unsafe or compromised robot policies. We demonstrate the effectiveness of our attacks against both state of the art action conditioned and text conditioned world models, showing a full end-to-end backdoor on a downstream DRL policy and a proof-of-concept for the VLA setting. Overall these findings necessitate research into more secure world models and reevaluating their position within the robot learning supply chain.


翻译:世界模型近年来在生成机器人训练数据或模拟真实环境方面,因其高效的数据利用能力而迅速普及且能力激增,许多研究提出将其整合到机器人学习流程中。尽管这些模型极具实用性,但本研究表明,世界模型为机器人学习供应链引入了一种独特且隐蔽的数据投毒入口点,可能导致在基于看似安全的真实训练数据完成训练后,部署不安全的或存在后门的机器人策略。与传统数据投毒技术直接向已出售或公开数据集植入危险轨迹不同,我们的新型攻击方法将恶意提示或有问题的转换动态注入到视觉安全的遥操作数据集中,这些恶意数据仅在被用作世界模型输入时才会激活。这可能导致生成合成型危险机器人训练轨迹,并最终形成不安全或受操控的机器人策略。我们针对当前最先进的动作条件型与文本条件型世界模型验证了攻击的有效性,展示了针对下游深度强化学习策略的全链路后门攻击,并在视觉-语言-动作(VLA)场景中完成了概念验证。总体而言,这些发现要求对更安全的世界模型展开研究,并重新评估其在机器人学习供应链中的位置。

0
下载
关闭预览

相关内容

【MIT博士论文】通过神经物理构建世界模型
专知会员服务
36+阅读 · 2025年4月3日
【UIUC博士论文】《从视频中进行机器人学习》
专知会员服务
25+阅读 · 2024年12月20日
面向机器学习模型安全的测试与修复
专知会员服务
55+阅读 · 2023年2月5日
专知会员服务
24+阅读 · 2021年8月22日
专知会员服务
49+阅读 · 2021年5月17日
一文读懂机器学习模型的选择与取舍
DBAplus社群
13+阅读 · 2019年8月25日
深度学习时代的图模型,清华发文综述图网络
GAN生成式对抗网络
13+阅读 · 2018年12月23日
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
52+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
【MIT博士论文】通过神经物理构建世界模型
专知会员服务
36+阅读 · 2025年4月3日
【UIUC博士论文】《从视频中进行机器人学习》
专知会员服务
25+阅读 · 2024年12月20日
面向机器学习模型安全的测试与修复
专知会员服务
55+阅读 · 2023年2月5日
专知会员服务
24+阅读 · 2021年8月22日
专知会员服务
49+阅读 · 2021年5月17日
相关基金
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
52+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员