We characterize the statistical efficiency of knowledge transfer through $n$ samples from a teacher to a probabilistic student classifier with input space $\mathcal S$ over labels $\mathcal A$. We show that privileged information at three progressive levels accelerates the transfer. At the first level, only samples with hard labels are known, via which the maximum likelihood estimator attains the minimax rate $\sqrt{{|{\mathcal S}||{\mathcal A}|}/{n}}$. The second level has the teacher probabilities of sampled labels available in addition, which turns out to boost the convergence rate lower bound to ${{|{\mathcal S}||{\mathcal A}|}/{n}}$. However, under this second data acquisition protocol, minimizing a naive adaptation of the cross-entropy loss results in an asymptotically biased student. We overcome this limitation and achieve the fundamental limit by using a novel empirical variant of the squared error logit loss. The third level further equips the student with the soft labels (complete logits) on ${\mathcal A}$ given every sampled input, thereby provably enables the student to enjoy a rate ${|{\mathcal S}|}/{n}$ free of $|{\mathcal A}|$. We find any Kullback-Leibler divergence minimizer to be optimal in the last case. Numerical simulations distinguish the four learners and corroborate our theory.
翻译:我们刻画了通过$n$个样本从教师到概率学生分类器(输入空间$\mathcal S$,标签集$\mathcal A$)进行知识迁移的统计效率。研究表明,三个递进层级的特权信息能加速迁移过程。第一层级仅提供硬标签样本,通过最大似然估计可达到极小极大速率$\sqrt{{|{\mathcal S}||{\mathcal A}|}/{n}}$。第二层级额外提供样本标签的教师概率,这使收敛速率下界提升至${{|{\mathcal S}||{\mathcal A}|}/{n}}$。然而在此数据获取协议下,直接最小化交叉熵损失会导致学生产生渐近偏差。我们通过引入新型经验均方误差对数几率损失克服该局限并达到基本极限。第三层级进一步向学生提供每个采样输入在$\mathcal A$上的软标签(完整对数几率),从而证明可使学生获得与$|{\mathcal A}|$无关的速率${|{\mathcal S}|}/{n}$。我们发现任何Kullback-Leibler散度最小化器在最后一种情况下均为最优。数值仿真区分了四种学习器并验证了理论结果。