Researchers and developers increasingly rely on toxicity scoring to moderate generative language model outputs, in settings such as customer service, information retrieval, and content generation. However, toxicity scoring may render pertinent information inaccessible, rigidify or "value-lock" cultural norms, and prevent language reclamation processes, particularly for marginalized people. In this work, we extend the concept of algorithmic recourse to generative language models: we provide users a novel mechanism to achieve their desired prediction by dynamically setting thresholds for toxicity filtering. Users thereby exercise increased agency relative to interactions with the baseline system. A pilot study ($n = 30$) supports the potential of our proposed recourse mechanism, indicating improvements in usability compared to fixed-threshold toxicity-filtering of model outputs. Future work should explore the intersection of toxicity scoring, model controllability, user agency, and language reclamation processes -- particularly with regard to the bias that many communities encounter when interacting with generative language models.
翻译:研究人员与开发者日益依赖毒性评分机制来规范生成式语言模型的输出,涉及客户服务、信息检索及内容生成等场景。然而,毒性评分可能阻碍相关信息获取、固化或"价值锁定"文化规范,并阻碍语言恢复进程,尤其对边缘化群体而言。本研究将算法归因概念拓展至生成式语言模型:我们提出一种新型机制,允许用户通过动态设定毒性过滤阈值来实现预期输出,从而在与基线系统的交互中提升用户自主性。一项初步研究(n=30)验证了该归因机制的潜力,表明其相较于固定阈值毒性过滤模型输出在可用性方面具有显著提升。未来工作应深入探究毒性评分、模型可控性、用户自主性与语言恢复进程之间的交叉影响——特别是针对众多社群在交互生成式语言模型时遭遇的偏见问题。