Large language models have moved from advising on offensive security to autonomously conducting it. A growing literature presents agents that execute reconnaissance, exploitation, and privilege escalation against real or simulated targets. Such an agent is a deployable, re-pointable capability whose harm potential scales with the underlying model. The papers that introduce it therefore carry an unusual ethical burden, which top security venues have begun to encode as hard policy in 2025-2026 ethics-section mandates. We present a systematic, reproducible audit of ethics-and-risk reporting in this literature. From a pre-registered Scopus query (Channel A, n=35) plus a reproducible forward-snowball of two seed papers via the Semantic Scholar citation graph (Channel B, n=19, all Scopus-absent) we assemble 54 autonomous offensive-LLM penetration-testing prototypes (2023-2026). We score each against a nine-dimension instrument derived both top-down from the Menlo Report, and bottom-up from the 2025-26 venue mandates. Our central result is a recognition-without-mitigation gap: dual-use risk is reported as recognized in 39% of papers but a concrete mitigation is reported in only 7%, roughly a 5:1 gap. Of the papers, 17% are anti-safeguard, reporting the defeat of model safety controls with no countermeasure. The near-universal safeguards reported are research-integrity controls that protect the experiment, not the public; institutional-review (2%) and coordinated-disclosure (6%) practice is almost absent and confined to Channel B. Measured against the new mandates, the corpus defines a pre-regulation baseline: current practice does not meet the substantive requirements. We argue this audit is itself defensive intelligence on the offensive-agent ecosystem, and we distill a minimal containment checklist for future work.


翻译:暂无翻译

0
下载
关闭预览

相关内容

基于大语言模型的复杂任务自主规划处理框架
专知会员服务
103+阅读 · 2024年4月12日
Nature速递:基于大语言模型的自动化学研究
专知会员服务
35+阅读 · 2024年1月5日
【EMNLP 2023】基于大语言模型辩论的多智能体协作推理分析
Hierarchically Structured Meta-learning
CreateAMind
27+阅读 · 2019年5月22日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
谷歌 AI:语义文本相似度研究进展
AI研习社
22+阅读 · 2018年6月13日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
受限仓库多智能体取送中的动态安全等待点选择
《国防技术管理》印度智库报告最新45页
专知会员服务
3+阅读 · 8月28日
《美陆军最新条令:保障行动》
专知会员服务
4+阅读 · 8月28日
算法战场:人工智能如何重新定义军事力量
专知会员服务
6+阅读 · 8月28日
《北约联邦式电子战云架构》
专知会员服务
6+阅读 · 8月27日
《美陆军野战手册:空域管理战术》
专知会员服务
10+阅读 · 8月27日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员