Privacy policies of websites are often lengthy and intricate. Privacy assistants assist in simplifying policies and making them more accessible and user friendly. The emergence of generative AI (genAI) offers new opportunities to build privacy assistants that can answer users questions about privacy policies. However, genAIs reliability is a concern due to its potential for producing inaccurate information. This study introduces GenAIPABench, a benchmark for evaluating Generative AI-based Privacy Assistants (GenAIPAs). GenAIPABench includes: 1) A set of questions about privacy policies and data protection regulations, with annotated answers for various organizations and regulations; 2) Metrics to assess the accuracy, relevance, and consistency of responses; and 3) A tool for generating prompts to introduce privacy documents and varied privacy questions to test system robustness. We evaluated three leading genAI systems ChatGPT-4, Bard, and Bing AI using GenAIPABench to gauge their effectiveness as GenAIPAs. Our results demonstrate significant promise in genAI capabilities in the privacy domain while also highlighting challenges in managing complex queries, ensuring consistency, and verifying source accuracy.
翻译:网站隐私政策通常冗长且复杂。隐私助手有助于简化政策,使其更易理解且用户友好。生成式AI(genAI)的出现为构建能够回答用户隐私政策相关问题的隐私助手提供了新机遇。然而,生成式AI的可靠性因其可能产生不准确信息而令人担忧。本研究提出GenAIPABench——一个用于评估基于生成式AI的隐私助手(GenAIPAs)的基准测试。GenAIPABench包含:1)一套关于隐私政策和数据保护法规的问题集,附带多种组织和法规的标注答案;2)评估响应准确性、相关性和一致性的指标;3)一种生成提示的工具,用于引入隐私文档和多样化隐私问题以测试系统鲁棒性。我们使用GenAIPABench评估了三个领先的生成式AI系统——ChatGPT-4、Bard和Bing AI,以衡量其作为GenAIPAs的有效性。结果表明,生成式AI在隐私领域展现出显著潜力,同时在处理复杂查询、确保一致性和验证信息来源准确性方面仍面临挑战。