In this paper, we develop a series of differential privacy (DP) algorithms from a family of random projections (RP) for general applications in machine learning, data mining, and information retrieval. Among the presented algorithms, iDP-SignRP is remarkably effective under the setting of ``individual differential privacy'' (iDP), based on sign random projections (SignRP). Also, DP-SignOPORP considerably improves existing algorithms in the literature under the standard DP setting, using ``one permutation + one random projection'' (OPORP), where OPORP is a variant of the celebrated count-sketch method with fixed-length binning and normalization. Without taking signs, among the DP-RP family, DP-OPORP achieves the best performance. Our key idea for improving DP-RP is to take only the signs, i.e., $sign(x_j) = sign\left(\sum_{i=1}^p u_i w_{ij}\right)$, of the projected data. The intuition is that the signs often remain unchanged when the original data ($u$) exhibit small changes (according to the ``neighbor'' definition in DP). In other words, the aggregation and quantization operations themselves provide good privacy protections. We develop a technique called ``smooth flipping probability'' that incorporates this intuitive privacy benefit of SignRPs and improves the standard DP bit flipping strategy. Based on this technique, we propose DP-SignOPORP which satisfies strict DP and outperforms other DP variants based on SignRP (and RP), especially when $\epsilon$ is not very large (e.g., $\epsilon = 5\sim10$). Moreover, if an application scenario accepts individual DP, then we immediately obtain an algorithm named iDP-SignRP which achieves excellent utilities even at small~$\epsilon$ (e.g., $\epsilon<0.5$).
翻译:本文针对机器学习、数据挖掘和信息检索领域的通用应用,从随机投影(RP)族中开发了一系列差分隐私(DP)算法。所提出的算法中,基于符号随机投影(SignRP)的iDP-SignRP在"个体差分隐私"(iDP)设定下效果显著。而DP-SignOPORP采用"一次排列+一次随机投影"(OPORP)方法,在标准DP设定下显著改进了现有文献算法——其中OPORP是经典count-sketch方法采用固定长度分箱与归一化处理的变体。在不采用符号的DP-RP族中,DP-OPORP取得了最佳性能。提升DP-RP的核心思路是仅取投影数据的符号:$sign(x_j) = sign\left(\sum_{i=1}^p u_i w_{ij}\right)$。其直觉依据在于,当原始数据($u$)发生微小变化时(符合DP中"邻近"定义),符号通常保持不变。换言之,聚合与量化操作本身即提供了良好的隐私保护。我们开发了名为"平滑翻转概率"的技术,该技术融合了SignRP的这一内在隐私优势,并改进了标准DP比特翻转策略。基于该技术,我们提出严格满足DP的DP-SignOPORP算法,其性能优于其他基于SignRP(及RP)的DP变体,尤其在$\epsilon$值并非极小的情况下(例如$\epsilon = 5\sim10$)。此外,若应用场景接受个体DP,则可直接获得名为iDP-SignRP的算法,该算法即使在较小$\epsilon$值下(例如$\epsilon<0.5$)仍能实现优异效用。