With the development of artificial intelligence, dialogue systems have been endowed with amazing chit-chat capabilities, and there is widespread interest and discussion about whether the generated contents are socially beneficial. In this paper, we present a new perspective of research scope towards building a safe, responsible, and modal dialogue system, including 1) abusive and toxic contents, 2) unfairness and discrimination, 3) ethics and morality issues, and 4) risk of misleading and privacy information. Besides, we review the mainstream methods for evaluating the safety of large models from the perspectives of exposure and detection of safety issues. The recent advances in methodologies for the safety improvement of both end-to-end dialogue systems and pipeline-based models are further introduced. Finally, we discussed six existing challenges towards responsible AI: explainable safety monitoring, continuous learning of safety issues, robustness against malicious attacks, multimodal information processing, unified research framework, and multidisciplinary theory integration. We hope this survey will inspire further research toward safer dialogue systems.
翻译:随着人工智能的发展,对话系统具备了令人惊叹的闲聊能力,其生成内容是否对社会有益引发了广泛关注与讨论。本文提出了构建安全、负责任且合乎道德的对话系统的研究视角,涵盖以下四大领域:1)辱骂与有毒内容;2)不公与歧视;3)伦理与道德问题;4)误导与隐私信息泄露风险。此外,我们从安全问题暴露与检测的角度,综述了评估大模型安全性的主流方法。进一步介绍了面向端到端对话系统与基于流水线模型的安全改进方法论的最新进展。最后,我们探讨了实现负责任人工智能面临的六大现有挑战:可解释的安全监控、安全问题的持续学习、对恶意攻击的鲁棒性、多模态信息处理、统一研究框架以及多学科理论整合。本综述旨在推动更安全的对话系统研究进一步发展。