An increasingly prevalent problem for intelligent technologies is text safety, as uncontrolled systems may generate recommendations to their users that lead to injury or life-threatening consequences. However, the degree of explicitness of a generated statement that can cause physical harm varies. In this paper, we distinguish types of text that can lead to physical harm and establish one particularly underexplored category: covertly unsafe text. Then, we further break down this category with respect to the system's information and discuss solutions to mitigate the generation of text in each of these subcategories. Ultimately, our work defines the problem of covertly unsafe language that causes physical harm and argues that this subtle yet dangerous issue needs to be prioritized by stakeholders and regulators. We highlight mitigation strategies to inspire future researchers to tackle this challenging problem and help improve safety within smart systems.
翻译:智能技术中日益普遍的问题是文本安全性,因为不受控制的系统可能向用户生成导致伤害或危及生命的建议。然而,可能造成身体伤害的生成语句其明确程度各不相同。本文首先区分了可能导致身体伤害的文本类型,并界定了一个特别未充分探索的类别:隐蔽不安全文本。随后,我们根据系统信息进一步细分该类别,并讨论缓解各子类别文本生成的解决方案。最终,我们的工作定义了导致身体伤害的隐蔽不安全语言问题,并主张这一微妙但危险的问题需要利益相关者和监管机构优先处理。我们强调缓解策略以激励未来研究者应对这一挑战性难题,并助力提升智能系统的安全性。