Concerns around future dangers from advanced AI often centre on systems hypothesised to have intrinsic characteristics such as agent-like behaviour, strategic awareness, and long-range planning. We label this cluster of characteristics as "Property X". Most present AI systems are low in "Property X"; however, in the absence of deliberate steering, current research directions may rapidly lead to the emergence of highly capable AI systems that are also high in "Property X". We argue that "Property X" characteristics are intrinsically dangerous, and when combined with greater capabilities will result in AI systems for which safety and control is difficult to guarantee. Drawing on several scholars' alternative frameworks for possible AI research trajectories, we argue that most of the proposed benefits of advanced AI can be obtained by systems designed to minimise this property. We then propose indicators and governance interventions to identify and limit the development of systems with risky "Property X" characteristics.
翻译:关于先进AI未来风险的担忧,通常聚焦于那些被假设具有内在特征的系统,如类主体行为、战略意识和长期规划。我们将这一特征集群标记为“属性X”。当前大多数AI系统的“属性X”水平较低;然而,若无刻意引导,当前的研究方向可能迅速催生出兼具高能力和高“属性X”特性的AI系统。我们认为“属性X”特征本质上是危险的,当与更强大的能力结合时,将导致难以保障安全性与可控性的AI系统。借鉴多位学者关于AI研究可能路径的替代框架,我们论证了先进AI的大部分预期收益可通过设计最小化该属性的系统来实现。最后,我们提出识别和限制具有风险性“属性X”特征的系统之指标与治理干预措施。