The private collection of multiple statistics from a population is a fundamental statistical problem. One possible approach to realize this is to rely on the local model of differential privacy (LDP). Numerous LDP protocols have been developed for the task of frequency estimation of single and multiple attributes. These studies mainly focused on improving the utility of the algorithms to ensure the server performs the estimations accurately. In this paper, we investigate privacy threats (re-identification and attribute inference attacks) against LDP protocols for multidimensional data following two state-of-the-art solutions for frequency estimation of multiple attributes. To broaden the scope of our study, we have also experimentally assessed five widely used LDP protocols, namely, generalized randomized response, optimal local hashing, subset selection, RAPPOR and optimal unary encoding. Finally, we also proposed a countermeasure that improves both utility and robustness against the identified threats. Our contributions can help practitioners aiming to collect users' statistics privately to decide which LDP mechanism best fits their needs.
翻译:从群体中私下收集多个统计量是一个基本的统计问题。实现这一目标的一种可能方法是依赖差分隐私的本地模型(LDP)。目前已开发出众多用于单一属性和多属性频率估计的LDP协议。这些研究主要侧重于提高算法的效用,以确保服务器能够准确执行估计。在本文中,我们研究了针对多维数据的LDP协议在隐私威胁(重识别和属性推断攻击)方面的问题,遵循了两种最新的多属性频率估计解决方案。为拓宽研究范围,我们还实验评估了五种广泛使用的LDP协议,即广义随机响应、最优局部哈希、子集选择、RAPPOR和最优一元编码。最后,我们提出了一种针对已识别威胁的防御措施,既提高了效用又增强了鲁棒性。我们的贡献能帮助旨在私下收集用户统计数据的从业者决定哪种LDP机制最符合其需求。