Through legislation and technical advances users gain more control over how their data is processed, and they expect online services to respect their privacy choices and preferences. However, data may be processed for many different purposes by several layers of algorithms that create complex data workflows. To date, there is no existing approach to automatically satisfy fine-grained privacy constraints of a user in a way which optimises the service provider's gains from processing. In this article, we propose a solution to this problem by modelling a data flow as a graph. User constraints and processing purposes are pairs of vertices which need to be disconnected in this graph. In general, this problem is NP-hard, thus, we propose several heuristics and algorithms. We discuss the optimality versus efficiency of our algorithms and evaluate them using synthetically generated data. On the practical side, our algorithms can provide nearly optimal solutions for tens of constraints and graphs of thousands of nodes, in a few seconds.
翻译:随着立法和技术进步,用户对其数据处理方式获得了更多控制权,并期望在线服务尊重其隐私选择与偏好。然而,数据可能经由多层算法为多种不同目的进行处理,形成复杂的数据工作流。迄今尚无现有方法能够自动满足用户细粒度的隐私约束,同时优化服务提供商从数据处理中获得的收益。本文提出了一种解决方案,通过将数据流建模为图结构,将用户约束与处理目的视为图中需要断开连接的顶点对。该问题通常属于NP难问题,因此我们提出了多种启发式算法。我们探讨了算法的最优性与效率之间的权衡,并使用合成生成数据对其进行评估。实践表明,我们的算法可在数秒内为包含数十个约束条件与数千个节点的图提供接近最优的解决方案。