The de-anonymization of users from anonymized microdata through matching or aligning with publicly-available correlated databases has been of scientific interest recently. While most of the rigorous analyses of database matching have focused on random-distortion models, the adversarial-distortion models have been wanting in the relevant literature. In this work, motivated by synchronization errors in the sampling of time-indexed microdata, matching (alignment) of random databases under adversarial column deletions is investigated. It is assumed that a constrained adversary, which observes the anonymized database, can delete up to a $\delta$ fraction of the columns (attributes) to hinder matching and preserve privacy. Column histograms of the two databases are utilized as permutation-invariant features to detect the column deletion pattern chosen by the adversary. The detection of the column deletion pattern is then followed by an exact row (user) matching scheme. The worst-case analysis of this two-phase scheme yields a sufficient condition for the successful matching of the two databases, under the near-perfect recovery condition. A more detailed investigation of the error probability leads to a tight necessary condition on the database growth rate, and in turn, to a single-letter characterization of the adversarial matching capacity. This adversarial matching capacity is shown to be significantly lower than the random matching capacity, where the column deletions occur randomly. Overall, our results analytically demonstrate the privacy-wise advantages of adversarial mechanisms over random ones during the publication of anonymized time-indexed data.
翻译:通过匹配或对齐公开可用的相关数据库来从匿名微数据中反匿名化用户,近年来引起了科学界的关注。尽管对数据库匹配的严格分析大多集中于随机失真模型,但相关文献中对抗性失真模型的研究相对不足。本文受时间索引微数据采样中的同步误差启发,研究了随机数据库在对抗性列删除下的匹配(对齐)问题。假设一个观察匿名数据库的受约束对抗者可以删除最多δ比例的列(属性)以阻碍匹配并保护隐私。利用两个数据库的列直方图作为置换不变特征来检测对抗者选择的列删除模式。检测列删除模式后,随后进行精确的行(用户)匹配方案。对该两阶段方案的最坏情况分析给出了在近乎完美恢复条件下成功匹配两个数据库的充分条件。对错误概率的更详细研究导出了数据库增长率的紧致必要条件,进而得到了对抗性匹配容量的单字母刻画。该对抗性匹配容量被证明显著低于随机匹配容量(后者列删除是随机的)。总体而言,我们的结果从分析角度证明了在发布匿名化时间索引数据时,对抗性机制相对于随机机制在隐私方面的优势。