In this work, we study practical heuristics to improve the performance of prefix-tree based algorithms for differentially private heavy hitter detection. Our model assumes each user has multiple data points and the goal is to learn as many of the most frequent data points as possible across all users' data with aggregate and local differential privacy. We propose an adaptive hyperparameter tuning algorithm that improves the performance of the algorithm while satisfying computational, communication and privacy constraints. We explore the impact of different data-selection schemes as well as the impact of introducing deny lists during multiple runs of the algorithm. We test these improvements using extensive experimentation on the Reddit dataset~\cite{caldas2018leaf} on the task of learning the most frequent words.
翻译:在本文中,我们研究了改进基于前缀树的差分隐私高频项检测算法性能的实用启发式方法。我们的模型假设每个用户拥有多个数据点,目标是在满足聚合差分隐私和局部差分隐私的前提下,尽可能多地识别所有用户数据中出现频率最高的数据点。我们提出了一种自适应超参数调优算法,在满足计算、通信和隐私约束的同时,提升了算法性能。我们探讨了不同数据选择方案的影响,以及在算法多次运行过程中引入拒绝列表所带来的影响。通过在Reddit数据集~\cite{caldas2018leaf}上针对学习最频繁词汇任务进行大量实验,我们验证了这些改进措施的有效性。