Differential privacy (DP) has the potential to enable privacy-preserving analysis on sensitive data, but requires analysts to judiciously spend a limited ``privacy loss budget'' $\epsilon$ across queries. Analysts conducting exploratory analyses do not, however, know all queries in advance and seldom have DP expertise. Thus, they are limited in their ability to specify $\epsilon$ allotments across queries prior to an analysis. To support analysts in spending $\epsilon$ efficiently, we propose a new interactive analysis paradigm, Measure-Observe-Remeasure, where analysts ``measure'' the database with a limited amount of $\epsilon$, observe estimates and their errors, and remeasure with more $\epsilon$ as needed. We instantiate the paradigm in an interactive visualization interface which allows analysts to spend increasing amounts of $\epsilon$ under a total budget. To observe how analysts interact with the Measure-Observe-Remeasure paradigm via the interface, we conduct a user study that compares the utility of $\epsilon$ allocations and findings from sensitive data participants make to the allocations and findings expected of a rational agent who faces the same decision task. We find that participants are able to use the workflow relatively successfully, including using budget allocation strategies that maximize over half of the available utility stemming from $\epsilon$ allocation. Their loss in performance relative to a rational agent appears to be driven more by their inability to access information and report it than to allocate $\epsilon$.
翻译:差分隐私(DP)有潜力实现对敏感数据的隐私保护分析,但要求分析者审慎地在查询间分配有限的"隐私损失预算"$\epsilon$。然而,进行探索性分析的分析者无法预先知晓所有查询,且通常缺乏DP专业知识。因此,他们在分析前指定各查询$\epsilon$分配方案的能力受限。为支持分析者高效使用$\epsilon$,我们提出一种新的交互式分析范式——测量-观察-再测量:分析者先用少量$\epsilon$对数据库进行"测量",观察估计值及其误差,再根据需要追加$\epsilon$进行"再测量"。我们将该范式实例化为交互式可视化界面,允许分析者在总预算范围内逐步增加$\epsilon$使用量。为观察分析者如何通过界面与测量-观察-再测量范式交互,我们开展用户研究,将参与者基于敏感数据作出的$\epsilon$分配方案及其发现结果,与面临相同决策任务的理性智能体预期方案进行效用比较。研究发现,参与者能相对成功地使用该工作流程,其预算分配策略可获取超过半数由$\epsilon$分配产生的可用效用。相较于理性智能体,参与者的性能损失主要源于信息获取与报告能力的不足,而非$\epsilon$分配能力缺陷。