Caution: This paper includes offensive words that could potentially cause unpleasantness. The fast-paced evolution of generative language models such as GPT-4 has demonstrated outstanding results in various NLP generation tasks. However, due to the potential generation of offensive words related to race or gender, various Controllable Text Generation (CTG) methods have been proposed to mitigate the occurrence of harmful words. However, existing CTG methods not only reduce toxicity but also negatively impact several aspects of the language model's generation performance, including topic consistency, grammar, and perplexity. This paper explores the limitations of previous methods and introduces a novel solution in the form of a simple Gated Toxicity Avoidance (GTA) that can be applied to any CTG method. We also evaluate the effectiveness of the proposed GTA by comparing it with state-of-the-art CTG methods across various datasets. Our findings reveal that gated toxicity avoidance efficiently achieves comparable levels of toxicity reduction to the original CTG methods while preserving the generation performance of the language model.
翻译:警告:本文包含可能引发不适的冒犯性词语。以GPT-4为代表的生成式语言模型在各类自然语言处理生成任务中展现出卓越性能,但其生成涉及种族或性别相关冒犯性词语的潜在风险催生了多种可控文本生成(CTG)方法以降低有害词出现概率。然而,现有CTG方法在降低毒性的同时,会对语言模型在主题连贯性、语法规范性和困惑度等生成性能维度产生负面影响。本文在剖析既有方法局限性的基础上,提出一种名为门控毒性规避(GTA)的简易解决方案,该方案可适配任意CTG方法。我们通过跨数据集对比实验,将所提GTA方法与当前最优CTG方法进行效果评估。结果表明,门控毒性规避能在保持语言模型生成性能的前提下,实现与原始CTG方法相当的毒性降低效果。