Rapid advances in Natural Language Processing (NLP) have revolutionized many fields, including healthcare. However, these advances raise significant privacy concerns, especially when pre-trained models fine-tuned and specialized on sensitive data can memorize and then expose and regurgitate personal information. This paper presents a privacy-preserving language modeling approach to address the problem of language models anonymization, and thus promote their sharing. Specifically, we propose both a Masking Language Modeling (MLM) methodology to specialize a BERT-like language model, and a Causal Language Modeling (CLM) methodology to specialize a GPT-like model that avoids the model from memorizing direct and indirect identifying information present in the training data. We have comprehensively evaluated our approaches using a medical dataset and compared them against different baselines. Our results indicate that by avoiding memorizing both direct and indirect identifiers during model specialization, our masking and causal language modeling schemes offer a good tradeoff for maintaining high privacy while retaining high utility.
翻译:自然语言处理(NLP)领域的快速发展已在包括医疗健康在内的多个领域引发革命性变革。然而,这些进步带来了显著的隐私担忧——尤其是当基于敏感数据微调与专精化的预训练模型可能记忆、泄露并复述个人信息时。本文提出一种隐私保护型语言建模方法,旨在解决语言模型匿名化问题,从而促进模型共享。具体而言,我们分别提出:用于专精化类BERT语言模型的掩码语言建模(MLM)方法,以及用于专精化类GPT模型的因果语言建模(CLM)方法,其能够避免模型记忆训练数据中存在的直接与间接标识信息。我们使用医疗数据集对上述方法进行了全面评估,并与多种基线方法进行了对比。研究表明,通过在模型专精化过程中避免记忆直接与间接标识符,所提出的掩码与因果语言建模方案能在维持高隐私保护水平的同时取得良好的效用平衡。