A prerequisite for safe autonomy-in-the-wild is safe testing-in-the-wild. Yet real-world autonomous tests face several unique safety challenges, both due to the possibility of causing harm during a test, as well as the risk of encountering new unsafe agent behavior through interactions with real-world and potentially malicious actors. We propose a framework for conducting safe autonomous agent tests on the open internet: agent actions are audited by a context-sensitive monitor that enforces a stringent safety boundary to stop an unsafe test, with suspect behavior ranked and logged to be examined by humans. We a design a basic safety monitor that is flexible enough to monitor existing LLM agents, and, using an adversarial simulated agent, we measure its ability to identify and stop unsafe situations. Then we apply the safety monitor on a battery of real-world tests of AutoGPT, and we identify several limitations and challenges that will face the creation of safe in-the-wild tests as autonomous agents grow more capable.
翻译:在野外实现安全自主性的先决条件是进行安全的野外测试。然而,真实世界的自主测试面临若干独特的安全挑战,既包括测试过程中可能造成伤害的风险,也包括通过与真实世界及潜在恶意行为者互动而遭遇新型不安全智能体行为的风险。我们提出一个框架,用于在开放互联网上进行安全自主智能体测试:智能体行为由上下文敏感监控器审计,该监控器执行严格安全边界以停止不安全测试,并将可疑行为分级记录以供人工审查。我们设计了一个基本的安全监控器,其灵活性足以监控现有的大语言模型智能体,并通过对抗性模拟智能体测量其识别和阻止不安全情境的能力。随后,我们将该安全监控器应用于AutoGPT的一系列真实世界测试,并识别出随着自主智能体能力增强,创建安全野外测试将面临的若干局限与挑战。