Artificial intelligence (AI) systems will increasingly be used to cause harm as they grow more capable. In fact, AI systems are already starting to be used to automate fraudulent activities, violate human rights, create harmful fake images, and identify dangerous toxins. To prevent some misuses of AI, we argue that targeted interventions on certain capabilities will be warranted. These restrictions may include controlling who can access certain types of AI models, what they can be used for, whether outputs are filtered or can be traced back to their user, and the resources needed to develop them. We also contend that some restrictions on non-AI capabilities needed to cause harm will be required. Though capability restrictions risk reducing use more than misuse (facing an unfavorable Misuse-Use Tradeoff), we argue that interventions on capabilities are warranted when other interventions are insufficient, the potential harm from misuse is high, and there are targeted ways to intervene on capabilities. We provide a taxonomy of interventions that can reduce AI misuse, focusing on the specific steps required for a misuse to cause harm (the Misuse Chain), and a framework to determine if an intervention is warranted. We apply this reasoning to three examples: predicting novel toxins, creating harmful images, and automating spear phishing campaigns.
翻译:人工智能系统在能力不断增强的同时,将越来越多地被用于造成伤害。事实上,AI系统已经开始被用于自动化欺诈活动、侵犯人权、制造有害虚假图像以及识别危险毒素。我们认为,为防止某些AI滥用行为,有必要对特定能力采取针对性干预措施。这些限制可能包括:控制谁可以访问特定类型的AI模型、模型的使用范围、输出内容是否需经过过滤或可追溯至用户,以及开发这些模型所需的资源。我们还主张,对造成伤害所需的非AI能力进行某些限制也是必要的。尽管能力限制存在减少实际用途多于减少滥用行为的风险(面临不利的“滥用-使用权衡”),但我们认为,当其他干预措施不足、滥用造成的潜在危害极高,且存在针对性的能力干预方式时,对能力进行干预是合理的。我们提出了减少AI滥用的干预措施分类体系,重点关注滥用行为造成伤害所需的具体步骤(滥用链),并构建了判断干预措施是否合理的评估框架。我们将这一理论应用于三个具体示例:预测新型毒素、制造有害图像以及自动化鱼叉式网络钓鱼攻击。