Metric Differential Privacy is a generalization of differential privacy tailored to address the unique challenges of text-to-text privatization. By adding noise to the representation of words in the geometric space of embeddings, words are replaced with words located in the proximity of the noisy representation. Since embeddings are trained based on word co-occurrences, this mechanism ensures that substitutions stem from a common semantic context. Without considering the grammatical category of words, however, this mechanism cannot guarantee that substitutions play similar syntactic roles. We analyze the capability of text-to-text privatization to preserve the grammatical category of words after substitution and find that surrogate texts consist almost exclusively of nouns. Lacking the capability to produce surrogate texts that correlate with the structure of the sensitive texts, we encompass our analysis by transforming the privatization step into a candidate selection problem in which substitutions are directed to words with matching grammatical properties. We demonstrate a substantial improvement in the performance of downstream tasks by up to $4.66\%$ while retaining comparative privacy guarantees.
翻译:度量差分隐私是一种针对文本到文本私有化独特挑战而量身定制的差分隐私泛化方法。通过在嵌入的几何空间中向单词表示添加噪声,单词会被替换为位于噪声表示邻近位置的单词。由于嵌入是基于单词共现训练的,该机制确保替换源于共同的语义上下文。然而,由于未考虑单词的语法类别,该机制无法保证替换扮演相似的句法角色。我们分析了文本到文本私有化在替换后保留单词语法类别的能力,发现替代文本几乎全部由名词构成。由于缺乏生成与敏感文本结构相关联的替代文本的能力,我们将私有化步骤转化为一个候选选择问题,其中替换被导向具有匹配语法属性的单词,从而整合了我们的分析。我们证明了下游任务性能的显著提升,最高达4.66%,同时保持了可比较的隐私保障。