The field of artificial intelligence (AI) has experienced remarkable progress in recent years, driven by the widespread adoption of open-source machine learning models in both research and industry. Considering the resource-intensive nature of training on vast datasets, many applications opt for models that have already been trained. Hence, a small number of key players undertake the responsibility of training and publicly releasing large pre-trained models, providing a crucial foundation for a wide range of applications. However, the adoption of these open-source models carries inherent privacy and security risks that are often overlooked. To provide a concrete example, an inconspicuous model may conceal hidden functionalities that, when triggered by specific input patterns, can manipulate the behavior of the system, such as instructing self-driving cars to ignore the presence of other vehicles. The implications of successful privacy and security attacks encompass a broad spectrum, ranging from relatively minor damage like service interruptions to highly alarming scenarios, including physical harm or the exposure of sensitive user data. In this work, we present a comprehensive overview of common privacy and security threats associated with the use of open-source models. By raising awareness of these dangers, we strive to promote the responsible and secure use of AI systems.
翻译:人工智能领域近年来取得了显著进展,这得益于开源机器学习模型在研究和工业界的广泛采用。考虑到在大型数据集上训练的资源密集型特性,许多应用选择使用已经训练好的模型。因此,少数关键参与方承担了训练并公开发布大型预训练模型的责任,为各类应用提供了关键基础。然而,采用这些开源模型会带来固有的隐私和安全风险,而这些风险往往被忽视。具体来说,一个看似不起眼的模型可能隐藏着隐藏功能,当被特定输入模式触发时,能够操纵系统的行为,例如指示自动驾驶汽车忽略其他车辆的存在。成功的隐私和安全攻击可能产生广泛影响,从服务中断等相对较小的损害,到涉及人身伤害或敏感用户数据泄露等令人高度担忧的场景。在本工作中,我们全面概述了与使用开源模型相关的常见隐私和安全威胁。通过提高对这些危险的认识,我们旨在促进AI系统的负责任和安全使用。