There has recently been an increased scientific interest in the de-anonymization of users in anonymized databases containing user-level microdata via multifarious matching strategies utilizing publicly available correlated data. Existing literature has either emphasized practical aspects where underlying data distribution is not required, with limited or no theoretical guarantees, or theoretical aspects with the assumption of complete availability of underlying distributions. In this work, we take a step towards reconciling these two lines of work by providing theoretical guarantees for the de-anonymization of random correlated databases without prior knowledge of data distribution. Motivated by time-indexed microdata, we consider database de-anonymization under both synchronization errors (column repetitions) and obfuscation (noise). By modifying the previously used replica detection algorithm to accommodate for the unknown underlying distribution, proposing a new seeded deletion detection algorithm, and employing statistical and information-theoretic tools, we derive sufficient conditions on the database growth rate for successful matching. Our findings demonstrate that a double-logarithmic seed size relative to row size ensures successful deletion detection. More importantly, we show that the derived sufficient conditions are the same as in the distribution-aware setting, negating any asymptotic loss of performance due to unknown underlying distributions.
翻译:近年来,利用公开关联数据通过多种匹配策略对包含用户级微观数据的匿名化数据库进行用户去匿名化的研究,引起了学术界的广泛关注。现有文献要么侧重实践层面,假设无需已知数据分布但缺乏或仅有有限的理论保证;要么侧重理论层面,假设数据分布完全已知。本文通过为随机关联数据库在无需数据分布先验的情况下提供去匿名化理论保证,旨在弥合这两类研究的鸿沟。受时间索引微观数据的启发,我们考虑同时存在同步错误(列重复)和混淆(噪声)条件下的数据库去匿名化。通过改进先前使用的副本检测算法以适应未知数据分布,提出新的种子式删除检测算法,并综合运用统计与信息论工具,推导出成功匹配所需的数据库增长率充分条件。结果表明,种子数量与行数呈双对数关系即可保证成功检测删除操作。更关键的是,我们证明所推导的充分条件与已知数据分布场景完全相同,从而消除了因未知数据分布导致的渐近性能损失。