Differential privacy is a widely adopted framework designed to safeguard the sensitive information of data providers within a data set. It is based on the application of controlled noise at the interface between the server that stores and processes the data, and the data consumers. Local differential privacy is a variant that allows data providers to apply the privatization mechanism themselves on their data individually. Therefore it provides protection also in contexts in which the server, or even the data collector, cannot be trusted. The introduction of noise, however, inevitably affects the utility of the data, particularly by distorting the correlations between individual data components. This distortion can prove detrimental to tasks such as causal discovery. In this paper, we consider various well-known locally differentially private mechanisms and compare the trade-off between the privacy they provide, and the accuracy of the causal structure produced by algorithms for causal learning when applied to data obfuscated by these mechanisms. Our analysis yields valuable insights for selecting appropriate local differentially private protocols for causal discovery tasks. We foresee that our findings will aid researchers and practitioners in conducting locally private causal discovery.
翻译:差分隐私是一种广泛采用的框架,旨在保护数据集中数据提供者的敏感信息。其核心在于:在存储和处理数据的服务器与数据消费者之间的接口处,施加受控噪声。局部差分隐私是其变体,允许数据提供者自行对其数据进行私有化处理,从而在服务器甚至数据收集者不可信的场景中也能提供保护。然而,噪声的引入不可避免地影响了数据的效用,尤其是扭曲了数据各分量之间的相关性。这种扭曲可能对因果发现等任务造成损害。本文研究了多种常见的局部差分隐私机制,比较了它们提供的隐私保护程度与基于这些机制混淆后的数据进行因果学习所得到的因果结构准确性之间的权衡。我们的分析为因果发现任务中如何选择合适的局部差分隐私协议提供了宝贵见解。预计我们的发现将有助于研究人员和从业者开展局部隐私下的因果发现工作。