The article explores the cultural shift from recording to deleting information in the digital age and its implications on privacy, intellectual property (IP), and Large Language Models like ChatGPT. It begins by defining a delete culture where information, in principle legal, is made unavailable or inaccessible because unacceptable or undesirable, especially but not only due to its potential to infringe on privacy or IP. Then it focuses on two strategies in this context: deleting, to make information unavailable; and blocking, to make it inaccessible. The article argues that both strategies have significant implications, particularly for machine learning (ML) models where information is not easily made unavailable. However, the emerging research area of Machine Unlearning (MU) is highlighted as a potential solution. MU, still in its infancy, seeks to remove specific data points from ML models, effectively making them 'forget' completely specific information. If successful, MU could provide a feasible means to manage the overabundance of information and ensure a better protection of privacy and IP. However, potential ethical risks, such as misuse, overuse, and underuse of MU, should be systematically studied to devise appropriate policies.
翻译:本文探讨了数字时代从记录信息到删除信息的文化转变,以及这种转变对隐私、知识产权以及ChatGPT等大型语言模型的影响。文章首先界定了“删除文化”这一概念:指原则上合法的信息因被认为不可接受或不受欢迎(尤其因其可能侵犯隐私或知识产权)而被使其不可获取或不可访问。随后重点阐述了此语境下的两种策略:删除——使信息不可获取;以及封锁——使信息不可访问。文章认为这两种策略均产生重大影响,特别是对信息难以被轻易清除的机器学习模型而言。不过,研究指出新兴领域“机器遗忘”作为潜在解决方案正受到关注。目前尚处萌芽阶段的机器遗忘技术致力于从机器学习模型中移除特定数据点,使其彻底“遗忘”指定信息。若该技术成熟,将为信息过载管理提供可行途径,并更有效地保护隐私与知识产权。然而,需系统研究其潜在的伦理风险(如滥用、过度使用及使用不足),以制定相应政策规范。