Modern database workloads are highly predictable: query streams are dominated by recurring jobs and templates, even when their arrival order is not known in advance. This motivates a learning-augmented view of online differentially private (DP) analytics: can algorithms utilize predictions about which queries will occur to improve utility under a single global privacy budget, while remaining robust when predictions are wrong? We study online DP query answering, where a curator must answer a stream $Q$ of $S$ linear queries arriving in uniformly random order under privacy budget $(ε,δ)$. We present LAPRAS, which assumes access to an oracle that outputs a prediction set of queries likely to appear in the stream and uses it to guide privacy spending. LAPRAS answers predicted queries using the offline-optimal Matrix Mechanism and answers the remaining queries online from a residual budget. To pace spending across an unknown number of unpredicted queries, we introduce Smooth Allocation, which forms an unbiased stopping-time estimate $\widehat{B}$ from the first $T=Θ(\log^2 S)$ unpredicted queries and continuously recalibrates per-query expenditure. Empirically, over two real datasets, we validate the intended consistency--robustness trade-off: LAPRAS achieves near-offline utility under high overlap and degrades gracefully to baseline-level performance when overlap is low.


翻译:摘要:现代数据库工作负载具有高度可预测性:查询流主要由重复性任务与模板构成,即使其到达顺序无法预先获知。这一特性催生了在线差分隐私分析的学习增强视角:算法能否利用关于未来查询类型的预测,在单一全局隐私预算约束下提升数据效用,同时在预测失准时保持鲁棒性?我们研究在线差分隐私查询应答问题,其中数据管理者需在隐私预算$(ε,δ)$约束下,对以均匀随机顺序到达的含有$S$个线性查询的流$Q$进行应答。本文提出LAPRAS方法,该方法假设可访问一个能输出流中可能出现查询的预测集的预言机,并据此指导隐私预算分配。LAPRAS对预测查询采用离线最优的矩阵机制进行应答,而对剩余查询则通过残差预算以在线方式处理。为应对未知数量的非预测查询的预算分配,我们提出平滑分配方法:该方法基于前$T=Θ(\log^2 S)$个非预测查询构造无偏停时估计量$\widehat{B}$,并持续动态校准每个查询的预算消耗。在两个真实数据集上的实验验证了预期的一致性与鲁棒性权衡效果:当查询重叠度高时,LAPRAS可实现接近离线状态的数据效用;当重叠度较低时,其性能优雅退化为基线水平。

0
下载
关闭预览

相关内容

专知会员服务
14+阅读 · 2021年9月14日
专知会员服务
41+阅读 · 2020年12月20日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
「联邦学习隐私保护 」最新2022研究综述
专知
16+阅读 · 2022年4月1日
搜索query意图识别的演进
DataFunTalk
13+阅读 · 2020年11月15日
联邦学习安全与隐私保护研究综述
专知
12+阅读 · 2020年8月7日
论文笔记之attention mechanism专题1:SA-Net(CVPR 2018)
统计学习与视觉计算组
16+阅读 · 2018年4月5日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
6+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员