The task of finding _Hierarchical_ Heavy Hitters (HHH) was introduced by Cormode et al. [VLDB 2003] as a generalisation of the heavy hitter problem. While finding HHH in data streams has been studied extensively, the question of releasing HHH when the underlying data is private remains unexplored. In this paper, we study differentially private HHH release in both the streaming and non-streaming setting. In the non-streaming setting, we show the surprising result that the relative error in estimating the residual count for any prefix is independent of the height of the hierarchy and the number of heavy hitters in the stream. Meanwhile, in the streaming setting, although the exact version of HHH has low global sensitivity (as counting queries are 1-sensitive), the approximation functions due to streaming have high global sensitivity, linear in the available space. Despite this obstacle, we show that the absolute error for estimating frequencies in the steaming setting is independent of the available space.
翻译:Cormode等人[VLDB 2003]提出的分层重击者(HHH)发现任务,是重击者问题的一种泛化形式。尽管数据流中的HHH发现已被广泛研究,但在底层数据涉及隐私时如何发布HHH仍是一个未探索的问题。本文研究了流式与非流式两种场景下差分隐私的HHH发布问题。在非流式场景中,我们得出一个令人惊讶的结论:任意前缀的残差计数估计的相对误差与层级高度及流中重击者数量无关。而在流式场景中,尽管HHH的精确版本具有低全局敏感度(因计数查询为1-敏感度),但流式处理导致的近似函数却具有与可用空间线性相关的高全局敏感度。尽管存在这一障碍,我们仍证明流式场景中频率估计的绝对误差与可用空间无关。