Recent advances in instruction-following large language models (LLMs) have led to dramatic improvements in a range of NLP tasks. Unfortunately, we find that the same improved capabilities amplify the dual-use risks for malicious purposes of these models. Dual-use is difficult to prevent as instruction-following capabilities now enable standard attacks from computer security. The capabilities of these instruction-following LLMs provide strong economic incentives for dual-use by malicious actors. In particular, we show that instruction-following LLMs can produce targeted malicious content, including hate speech and scams, bypassing in-the-wild defenses implemented by LLM API vendors. Our analysis shows that this content can be generated economically and at cost likely lower than with human effort alone. Together, our findings suggest that LLMs will increasingly attract more sophisticated adversaries and attacks, and addressing these attacks may require new approaches to mitigations.
翻译:近年来,遵循指令的大型语言模型(LLMs)的进步显著提升了多项自然语言处理任务的性能。不幸的是,我们发现这种增强的能力放大了这些模型被恶意利用的双重用途风险。由于指令遵循能力使得计算机安全领域的标准攻击成为可能,双重用途问题难以防范。这些指令遵循型LLMs的能力为恶意行为者提供了强大的经济激励以实现双重用途。具体而言,我们证明指令遵循型LLMs能够生成针对性的恶意内容,包括仇恨言论和诈骗,绕过LLM API供应商部署的现网防御措施。我们的分析表明,这类内容可以以低于人类单独努力的成本经济地生成。综合来看,我们的研究结果暗示LLMs将日益吸引更复杂的对手和攻击,而应对这些攻击可能需要新的缓解方法。