Document understanding models are increasingly employed by companies to supplant humans in processing sensitive documents, such as invoices, tax notices, or even ID cards. However, the robustness of such models to privacy attacks remains vastly unexplored. This paper presents CDMI, the first reconstruction attack designed to extract sensitive fields from the training data of these models. We attack LayoutLM and BROS architectures, demonstrating that an adversary can perfectly reconstruct up to 4.1% of the fields of the documents used for fine-tuning, including some names, dates, and invoice amounts up to six-digit numbers. When our reconstruction attack is combined with a membership inference attack, our attack accuracy escalates to 22.5%. In addition, we introduce two new end-to-end metrics and evaluate our approach under various conditions: unimodal or bimodal data, LayoutLM or BROS backbones, four fine-tuning tasks, and two public datasets (FUNSD and SROIE). We also investigate the interplay between overfitting, predictive performance, and susceptibility to our attack. We conclude with a discussion on possible defenses against our attack and potential future research directions to construct robust document understanding models.
翻译:文档理解模型正日益被企业用于替代人类处理敏感文档,例如发票、税务通知甚至身份证。然而,此类模型对隐私攻击的鲁棒性仍鲜有研究。本文提出CDMI,这是首个旨在从这些模型的训练数据中提取敏感字段的重建攻击方法。我们对LayoutLM和BROS架构进行攻击,证明攻击者可以完美重建最多4.1%用于微调的文档字段,包括部分姓名、日期以及最多六位数的发票金额。当我们的重建攻击与成员推断攻击相结合时,攻击准确率提升至22.5%。此外,我们引入了两种新的端到端指标,并在多种条件下评估了我们的方法:单模态或双模态数据、LayoutLM或BROS骨干网络、四种微调任务以及两个公开数据集(FUNSD和SROIE)。我们还研究了过拟合、预测性能与攻击易感性之间的相互影响。最后,我们讨论了针对该攻击的可能防御措施以及构建鲁棒文档理解模型的潜在未来研究方向。