Language Models (LMs) have been shown to leak information about training data through sentence-level membership inference and reconstruction attacks. Understanding the risk of LMs leaking Personally Identifiable Information (PII) has received less attention, which can be attributed to the false assumption that dataset curation techniques such as scrubbing are sufficient to prevent PII leakage. Scrubbing techniques reduce but do not prevent the risk of PII leakage: in practice scrubbing is imperfect and must balance the trade-off between minimizing disclosure and preserving the utility of the dataset. On the other hand, it is unclear to which extent algorithmic defenses such as differential privacy, designed to guarantee sentence- or user-level privacy, prevent PII disclosure. In this work, we introduce rigorous game-based definitions for three types of PII leakage via black-box extraction, inference, and reconstruction attacks with only API access to an LM. We empirically evaluate the attacks against GPT-2 models fine-tuned with and without defenses on three domains: case law, health care, and e-mails. Our main contributions are (i) novel attacks that can extract up to 10$\times$ more PII sequences than existing attacks, (ii) showing that sentence-level differential privacy reduces the risk of PII disclosure but still leaks about 3% of PII sequences, and (iii) a subtle connection between record-level membership inference and PII reconstruction.
翻译:语言模型(LMs)已被证明会通过句子级别的成员推理和重建攻击泄露训练数据中的信息。关于LMs泄露个人可识别信息(PII)风险的研究较少,这归因于一个错误假设:诸如数据清洗等数据集处理技术足以防止PII泄露。清洗技术能降低但无法消除PII泄露风险:实践中清洗并不完美,必须在最小化泄露与保留数据集效用之间权衡。另一方面,诸如差分隐私等旨在保障句子级别或用户级别隐私的算法防御手段能在多大程度上防止PII泄露尚不明确。本文针对仅通过API访问语言模型的三种PII泄露类型(黑盒提取、推理和重建攻击),引入了严格的基于博弈的定义。我们针对在三个领域(判例法、医疗保健和电子邮件)上使用或不使用防御手段微调的GPT-2模型进行实证攻击评估。主要贡献包括:(i)提出新型攻击方法,其提取的PII序列数量比现有攻击多出10倍以上;(ii)证明句子级别差分隐私虽能降低PII泄露风险,但仍有约3%的PII序列被泄露;(iii)发现记录级成员推理与PII重建之间存在微妙关联。