The Collaborative Research Cycle (CRC) is a National Institute of Standards and Technology (NIST) benchmarking program intended to strengthen understanding of tabular data deidentification technologies. Deidentification algorithms are vulnerable to the same bias and privacy issues that impact other data analytics and machine learning applications, and can even amplify those issues by contaminating downstream applications. This paper summarizes four CRC contributions: theoretical work on the relationship between diverse populations and challenges for equitable deidentification; public benchmark data focused on diverse populations and challenging features; a comprehensive open source suite of evaluation metrology for deidentified datasets; and an archive of more than 450 deidentified data samples from a broad range of techniques. The initial set of evaluation results demonstrate the value of these tools for investigations in this field.
翻译:协作研究周期(CRC)是美国国家标准与技术研究院(NIST)的基准测试项目,旨在加深对表格数据去标识化技术的理解。去标识化算法容易受到影响其他数据分析与机器学习应用的相同偏差和隐私问题的影响,甚至可能通过污染下游应用而放大这些问题。本文总结了CRC的四项贡献:关于多样化群体与公平去标识化挑战之间关系的理论工作;聚焦多样化群体及挑战性特征的公开基准数据;一套全面的开源去标识化数据集评估计量体系;以及包含来自广泛技术的450多个去标识化数据样本的档案。初步评估结果证明了这些工具在该领域研究中的价值。