In the rapidly evolving landscape of artificial intelligence, multimodal learning systems (MMLS) have gained traction for their ability to process and integrate information from diverse modality inputs. Their expanding use in vital sectors such as healthcare has made safety assurance a critical concern. However, the absence of systematic research into their safety is a significant barrier to progress in this field. To bridge the gap, we present the first taxonomy that systematically categorizes and assesses MMLS safety. This taxonomy is structured around four fundamental pillars that are critical to ensuring the safety of MMLS: robustness, alignment, monitoring, and controllability. Leveraging this taxonomy, we review existing methodologies, benchmarks, and the current state of research, while also pinpointing the principal limitations and gaps in knowledge. Finally, we discuss unique challenges in MMLS safety. In illuminating these challenges, we aim to pave the way for future research, proposing potential directions that could lead to significant advancements in the safety protocols of MMLS.
翻译:在人工智能快速发展的背景下,多模态学习系统因其处理与整合多模态输入信息的能力而备受关注。随着其在医疗等关键领域的广泛应用,安全保障已成为亟需关注的核心问题。然而,目前该领域尚未形成系统性的安全研究,这成为制约技术发展的重大瓶颈。为填补这一空白,我们首次提出系统化分类与评估多模态学习系统安全的分类体系。该分类体系围绕确保系统安全的四大核心支柱构建:鲁棒性、对齐性、可监控性与可操控性。基于此分类体系,我们系统综述了现有方法、基准测试及研究现状,并指明了当前存在的主要技术局限与认知空白。最后,我们探讨了多模态学习系统安全面临的独特挑战。通过揭示这些挑战,本文旨在为未来研究指明方向,提出可能推动多模态学习系统安全协议取得重大突破的潜在研究路径。