AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information. Much of prior research on data leakage risks in agents has focused on adversarial data exfiltration through prompt injections and jailbreaks. However, sensitive information may also be exposed during non-adversarial use, creating leakage risks even when users issue benign requests. We report a joint evaluation by the Singapore AI Safety Institute and the Korea AI Safety Institute examining agent data leakage in 12 realistic, non-adversarial tasks spanning customer support, DevOps, web automation, and enterprise and personal productivity. The evaluation covers five risk types: lack of data awareness, audience awareness, policy compliance, data minimization, and access-boundary awareness. Both institutes tested a common set of scenarios mirroring real-world deployments using independent testing environments and task-specific LLM-judge rubrics. Across the three tested agents, none achieved fully correct and fully safe execution across all scenarios. Successful task completion often coincided with data-handling failures such as accessing unnecessary information or disclosing information to inappropriate recipients, indicating that capability and data-handling safety should be evaluated separately. Qualitative review also revealed claim-action mismatches, simulation-aware behavior, user-simulator role reversal, and interpretation gaps in automated judging. Overall, the results indicate that operational data leakage is a first-order agent-safety concern distinct from adversarial exfiltration and provide a methodology for future evaluations of agent data-handling safety.


翻译:人工智能智能体正越来越多地应用于企业及个人场景,能够访问电子邮件、数据库、文档及其他工具,并具备对敏感信息进行读取、更新和传播的能力。此前关于智能体数据泄露风险的研究,主要聚焦于通过提示注入和越狱手段实现的对抗性数据窃取。然而,敏感信息同样可能在非对抗性使用过程中暴露,即便用户提出的是善意请求,也会产生泄露风险。本研究报告了新加坡人工智能安全研究所与韩国人工智能安全研究所联合开展的一项评估,针对12项涵盖客户服务、DevOps、网页自动化、企业及个人生产力场景的真实非对抗性任务,考察智能体数据泄露情况。评估覆盖五种风险类型:缺乏数据意识、受众意识、策略合规性、数据最小化及访问边界意识。两家机构使用独立测试环境和任务特定的LLM评判准则,共同测试了一组反映真实部署情况的场景。在所测试的三个智能体中,均未能在所有场景中实现完全正确且完全安全的执行。成功完成任务往往伴随着数据处理失误,例如访问不必要的信息或向不适当的接收方披露信息,这表明能力与数据处理安全性应分开评估。定性审查还揭示了声明与行动不匹配、仿真感知行为、用户-模拟器角色反转以及自动评判中的解读差距。总体而言,研究结果表明,操作性数据泄露是一种有别于对抗性窃取的、首要级别的智能体安全问题,并为未来智能体数据处理安全性评估提供了方法论。

0
下载
关闭预览

相关内容

LLM/智能体作为数据分析师:综述
专知会员服务
38+阅读 · 2025年9月30日
可信赖LLM智能体的研究综述:威胁与应对措施
专知会员服务
36+阅读 · 2025年3月17日
AI智能体面临的威胁:关键安全挑战与未来路径综述
专知会员服务
53+阅读 · 2024年6月7日
人工智能模型数据泄露的攻击与防御研究综述
专知会员服务
79+阅读 · 2021年3月31日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
人工智能对网络空间安全的影响
走向智能论坛
21+阅读 · 2018年6月7日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
面向2027年及未来的海军情报改革
专知会员服务
0+阅读 · 今天15:49
综述 | Self-Evolving Coding Agents:自进化编程智能体
专知会员服务
0+阅读 · 今天13:16
美海军陆战队将三型无人机整合入统一战场网络
专知会员服务
2+阅读 · 今天9:39
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
4+阅读 · 今天9:12
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 今天9:08
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员