In many real-world applications, in particular due to recent developments in the privacy landscape, training data may be aggregated to preserve the privacy of sensitive training labels. In the learning from label proportions (LLP) framework, the dataset is partitioned into bags of feature-vectors which are available only with the sum of the labels per bag. A further restriction, which we call learning from bag aggregates (LBA) is where instead of individual feature-vectors, only the (possibly weighted) sum of the feature-vectors per bag is available. We study whether such aggregation techniques can provide privacy guarantees under the notion of label differential privacy (label-DP) previously studied in for e.g. [Chaudhuri-Hsu'11, Ghazi et al.'21, Esfandiari et al.'22]. It is easily seen that naive LBA and LLP do not provide label-DP. Our main result however, shows that weighted LBA using iid Gaussian weights with $m$ randomly sampled disjoint $k$-sized bags is in fact $(\varepsilon, \delta)$-label-DP for any $\varepsilon > 0$ with $\delta \approx \exp(-\Omega(\sqrt{k}))$ assuming a lower bound on the linear-mse regression loss. Further, this preserves the optimum over linear mse-regressors of bounded norm to within $(1 \pm o(1))$-factor w.p. $\approx 1 - \exp(-\Omega(m))$. We emphasize that no additive label noise is required. The analogous weighted-LLP does not however admit label-DP. Nevertheless, we show that if additive $N(0, 1)$ noise can be added to any constant fraction of the instance labels, then the noisy weighted-LLP admits similar label-DP guarantees without assumptions on the dataset, while preserving the utility of Lipschitz-bounded neural mse-regression tasks. Our work is the first to demonstrate that label-DP can be achieved by randomly weighted aggregation for regression tasks, using no or little additive noise.
翻译:在许多实际应用中,特别是由于隐私领域的最新发展,训练数据可能会被聚合以保护敏感训练标签的隐私。在从标签比例学习(LLP)框架中,数据集被划分为特征向量的包,每个包仅提供标签总和。另一种更严格的限制称为从包聚合学习(LBA),其中每个包仅提供(可能加权的)特征向量总和,而非单个特征向量。我们研究此类聚合技术能否在标签差分隐私(label-DP)概念下提供隐私保障。此前已有研究关注标签差分隐私,例如[Chaudhuri-Hsu'11, Ghazi等人'21, Esfandiari等人'22]。显而易见,朴素LBA和LLP并不能提供标签差分隐私。然而,我们的主要结果表明,使用独立同分布的高斯权重,且包含m个随机抽样的不相交k大小包,在假设线性均方回归损失存在下界时,实际上对于任意ε>0且δ≈exp(-Ω(√k))的情况,加权LBA可实现(ε, δ)-标签差分隐私。此外,该方法能以约1-exp(-Ω(m))的概率将有界范数线性均方回归器的最优值保持在(1±o(1))因子内。需要强调的是,无需添加标签噪声。但类似的加权LLP则无法提供标签差分隐私。尽管如此,我们证明如果能够对任意恒定比例的实例标签添加N(0,1)噪声,则带噪声的加权LLP可在无需数据集假设的情况下提供类似的标签差分隐私保障,同时保持利普希茨有界神经网络均方回归任务的实用性。我们的工作是首个证明通过随机加权聚合可在回归任务中实现标签差分隐私(无需或仅需少量加性噪声)的研究。