Concerns about data privacy are omnipresent, given the increasing usage of digital applications and their underlying business model that includes selling user data. Location data is particularly sensitive since they allow us to infer activity patterns and interests of users, e.g., by categorizing visited locations based on nearby points of interest (POI). On top of that, machine learning methods provide new powerful tools to interpret big data. In light of these considerations, we raise the following question: What is the actual risk that realistic, machine learning based privacy attacks can obtain meaningful semantic information from raw location data, subject to inaccuracies in the data? In response, we present a systematic analysis of two attack scenarios, namely location categorization and user profiling. Experiments on the Foursquare dataset and tracking data demonstrate the potential for abuse of high-quality spatial information, leading to a significant privacy loss even with location inaccuracy of up to 200m. With location obfuscation of more than 1 km, spatial information hardly adds any value, but a high privacy risk solely from temporal information remains. The availability of public context data such as POIs plays a key role in inference based on spatial information. Our findings point out the risks of ever-growing databases of tracking data and spatial context data, which policymakers should consider for privacy regulations, and which could guide individuals in their personal location protection measures.
翻译:数据隐私问题无处不在,随着数字应用的日益普及及其底层商业模式(包括出售用户数据),这一问题愈发凸显。位置数据尤为敏感,因为它能够推断用户的活动模式和兴趣,例如,通过根据附近兴趣点(POI)对访问位置进行分类。此外,机器学习方法为解读大数据提供了新的强大工具。基于这些考量,我们提出以下问题:在数据存在不准确性的情况下,基于机器学习的现实隐私攻击能从原始位置数据中获取有意义的语义信息的实际风险有多大?为此,我们对两种攻击场景——位置分类和用户画像——进行了系统分析。在Foursquare数据集和跟踪数据上的实验表明,高质量空间信息存在被滥用的可能,即使在位置误差高达200米的情况下,也会导致显著的隐私损失。当位置混淆超过1公里时,空间信息几乎无法增加任何价值,但仅凭时间信息仍存在较高的隐私风险。诸如POI等公共上下文数据的可用性在基于空间信息的推断中起着关键作用。我们的研究结果揭示了日益增长的跟踪数据和空间上下文数据库所蕴含的风险,政策制定者应将其纳入隐私法规考量,同时这些结果也可指导个人采取位置保护措施。