The widespread exchange of digital documents in various domains has resulted in abundant private information being shared. This proliferation necessitates redaction techniques to protect sensitive content and user privacy. While numerous redaction methods exist, their effectiveness varies, with some proving more robust than others. As such, the literature proposes several deanonymization techniques, raising awareness of potential privacy threats. However, while none of these methods are successful against the most effective redaction techniques, these attacks only focus on the anonymized tokens and ignore the sentence context. In this paper, we propose RedactBuster, the first deanonymization model using sentence context to perform Named Entity Recognition on reacted text. Our methodology leverages fine-tuned state-of-the-art Transformers and Deep Learning models to determine the anonymized entity types in a document. We test RedactBuster against the most effective redaction technique and evaluate it using the publicly available Text Anonymization Benchmark (TAB). Our results show accuracy values up to 0.985 regardless of the document nature or entity type. In raising awareness of this privacy issue, we propose a countermeasure we call character evasion that helps strengthen the secrecy of sensitive information. Furthermore, we make our model and testbed open-source to aid researchers and practitioners in evaluating the resilience of novel redaction techniques and enhancing document privacy.
翻译:数字文档在各领域的广泛交换导致大量私人信息被共享。这种扩散现象需要采用编辑技术来保护敏感内容和用户隐私。尽管存在多种编辑方法,其有效性各有差异,部分方法被证明比其他方法更稳健。为此,文献中提出了多种去匿名化技术,旨在提高人们对潜在隐私威胁的认识。然而,虽然这些方法均无法攻破最有效的编辑技术,但这些攻击仅关注匿名化标记而忽略了句子上下文。本文提出RedactBuster——首个利用句子上下文对已编辑文本执行命名实体识别的去匿名化模型。我们的方法通过微调最先进的Transformer和深度学习模型,确定文档中匿名化实体的类型。我们针对最有效的编辑技术测试了RedactBuster,并使用公开的文本匿名化基准(TAB)进行评估。结果表明,无论文档性质或实体类型如何,准确率最高可达0.985。为引起对这一隐私问题的重视,我们提出一种名为字符规避的对抗措施,有助于强化敏感信息的保密性。此外,我们将模型和测试平台开源,以帮助研究人员和从业者评估新型编辑技术的鲁棒性并增强文档隐私保护。