This paper considers the problem of the private release of sample means of speed values from traffic datasets. Our key contribution is the development of user-level differentially private algorithms that incorporate carefully chosen parameter values to ensure low estimation errors on real-world datasets, while ensuring privacy. We test our algorithms on ITMS (Intelligent Traffic Management System) data from an Indian city, where the speeds of different buses are drawn in a potentially non-i.i.d. manner from an unknown distribution, and where the number of speed samples contributed by different buses is potentially different. We then apply our algorithms to a synthetic dataset, generated based on the ITMS data, having either a large number of users or a large number of samples per user. Here, we provide recommendations for the choices of parameters and algorithm subroutines that result in low estimation errors, while guaranteeing user-level privacy.
翻译:本文考虑了交通数据集中速度值样本均值的隐私发布问题。我们的核心贡献在于开发了用户级差分隐私算法,该算法通过精心选取参数值,在确保隐私的同时,能对真实世界数据集实现较低的估计误差。我们在印度某城市的ITMS(智能交通管理系统)数据上测试了所提算法——该数据中不同公交车的速度样本可能以非独立同分布方式抽取自未知分布,且不同车辆贡献的速度样本数量可能存在差异。随后,我们将算法应用于基于ITMS数据生成的合成数据集,该数据集包含大量用户或每个用户拥有大量样本。在此,我们提供了参数选择与算法子程序的建议,以在保证用户级隐私的前提下实现较低的估计误差。