Software applications have become an omnipresent part of modern society. The consequent privacy policies of these applications play a significant role in informing customers how their personal information is collected, stored, and used. However, customers rarely read and often fail to understand privacy policies because of the ``Privacy Policy Reading Phobia'' (PPRP). To tackle this emerging challenge, we propose the first framework that can automatically generate privacy nutrition labels from privacy policies. Based on our ground truth applications about the Data Safety Report from the Google Play app store, our framework achieves a 0.75 F1-score on generating first-party data collection practices and an average of 0.93 F1-score on general security practices. We also analyse the inconsistencies between ground truth and curated privacy nutrition labels on the market, and our framework can detect 90.1% under-claim issues. Our framework demonstrates decent generalizability across different privacy nutrition label formats, such as Google's Data Safety Report and Apple's App Privacy Details.
翻译:软件应用已成为现代社会中无处不在的一部分。随之而来的隐私政策在告知客户其个人信息如何被收集、存储和使用方面发挥着重要作用。然而,由于 “隐私政策阅读恐惧症”(PPRP),客户很少阅读且常常难以理解隐私政策。为应对这一新兴挑战,我们提出了首个能够从隐私政策自动生成隐私营养标签的框架。基于我们关于 Google Play 应用商店中数据安全报告的真实应用数据,该框架在生成第一方数据收集实践方面达到了 0.75 的 F1 分数,在通用安全实践方面平均 F1 分数为 0.93。我们还分析了市场中真实数据与精心设计的隐私营养标签之间的不一致性,该框架能够检测出 90.1% 的低报问题。我们的框架在多种隐私营养标签格式(如 Google 的数据安全报告和 Apple 的应用隐私详情)中展现出良好的泛化能力。