Large language models (LLMs) have exhibited impressive capabilities in comprehending complex instructions. However, their blind adherence to provided instructions has led to concerns regarding risks of malicious use. Existing defence mechanisms, such as model fine-tuning or output censorship using LLMs, have proven to be fallible, as LLMs can still generate problematic responses. Commonly employed censorship approaches treat the issue as a machine learning problem and rely on another LM to detect undesirable content in LLM outputs. In this paper, we present the theoretical limitations of such semantic censorship approaches. Specifically, we demonstrate that semantic censorship can be perceived as an undecidable problem, highlighting the inherent challenges in censorship that arise due to LLMs' programmatic and instruction-following capabilities. Furthermore, we argue that the challenges extend beyond semantic censorship, as knowledgeable attackers can reconstruct impermissible outputs from a collection of permissible ones. As a result, we propose that the problem of censorship needs to be reevaluated; it should be treated as a security problem which warrants the adaptation of security-based approaches to mitigate potential risks.
翻译:大型语言模型在理解复杂指令方面展现了令人印象深刻的能力。然而,它们对给定指令的盲目遵循引发了对其被恶意利用风险的担忧。现有的防御机制,如模型微调或使用LLM进行输出审查,已被证明存在缺陷,因为LLM仍能生成有问题的回复。常用的审查方法将这一问题视为机器学习问题,并依赖另一个语言模型来检测LLM输出中的不良内容。在本文中,我们揭示了这种语义审查方法的理论局限性。具体而言,我们证明语义审查可被视为一个不可判定问题,突显了因LLM的程序化和指令遵循能力而产生的内在审查挑战。此外,我们认为这些挑战超越了语义审查范畴,因为具备知识的攻击者可以从一系列允许的输出中重构出被禁止的内容。因此,我们提出需要重新审视审查问题;它应被视为一个安全问题,需要采用基于安全的方法来缓解潜在风险。