We study the problem of multiclass PAC learning with bandit feedback in the realizable setting. In this framework, there is an unknown data distribution over an instance space $\mathcal{X}$ and a label space $\mathcal{Y}$, as in classical multiclass PAC learning, but the learner does not observe the labels of the i.i.d. training examples. Instead, in each round, it receives an unlabeled instance, predicts its label, and receives bandit feedback indicating only whether the prediction is correct. Despite this restriction, the goal remains the same as in classical PAC learning. We provide a general characterization of the optimal sample complexity of this problem, sharp for every concept class up to logarithmic factors. Our characterization is based on a new combinatorial dimension, termed the bandit $\mathrm{DS}$ dimension, defined via generalized combinatorial structures we call pseudo-boxes. These extend the pseudo-cubes underlying the $\mathrm{DS}$ dimension by allowing a different number of neighbors in each coordinate. In contrast to the $\mathrm{DS}$ dimension, which governs the full-information setting by counting the number of coordinates in the pseudo-cube, the bandit $\mathrm{DS}$ dimension aggregates the number of neighbors across coordinates, leading to a characterization in which the sample complexity scales with the total number of neighbors. We also propose a general learning algorithm achieving the upper bound, based on an algorithmic principle called ListCascade, which connects bandit learning to list learning and may be of independent interest.


翻译:我们研究在可实现设置下带有bandit反馈的多类PAC学习问题。在该框架中,存在实例空间$\mathcal{X}$和标签空间$\mathcal{Y}$上的未知数据分布(与经典多类PAC学习相同),但学习器无法观测到独立同分布训练样本的标签。相反,在每一轮中,它接收一个未标记实例,预测其标签,并仅接收到指示预测是否正确的bandit反馈。尽管存在这一限制,但目标仍与经典PAC学习相同。我们针对该问题的最优样本复杂度给出了普适性刻画,该刻画对于每个概念类在多项式因子内均达到最优。我们的刻画基于一个新的组合维度——称为bandit $\mathrm{DS}$维度——它通过我们称之为伪盒的广义组合结构进行定义。这些结构扩展了构成$\mathrm{DS}$维度基础的伪立方体,允许每个坐标具有不同数量的邻域。与通过统计伪立方体中坐标数量来刻画全信息设置的$\mathrm{DS}$维度不同,bandit $\mathrm{DS}$维度聚合了各坐标的邻域数量,使得样本复杂度与邻域总数成比例的刻画得以实现。我们还提出了一种名为ListCascade的通用学习算法来实现上界,该算法将bandit学习与列表学习相连接,可能具有独立的研究价值。

0
下载
关闭预览

相关内容

【MIT】反偏差对比学习,Debiased Contrastive Learning
专知会员服务
92+阅读 · 2020年7月4日
基于深度神经网络的少样本学习综述
专知会员服务
173+阅读 · 2020年4月22日
小样本学习(Few-shot Learning)综述
机器之心
18+阅读 · 2019年4月1日
半监督深度学习小结:类协同训练和一致性正则化
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
14+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
1+阅读 · 今天14:36
《战争中的大语言模型监管》
专知会员服务
2+阅读 · 今天14:26
边缘计算的军事应用
专知会员服务
8+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
相关VIP内容
【MIT】反偏差对比学习,Debiased Contrastive Learning
专知会员服务
92+阅读 · 2020年7月4日
基于深度神经网络的少样本学习综述
专知会员服务
173+阅读 · 2020年4月22日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
14+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员