Black-Box Knowledge Distillation (B2KD) is a formulated problem for cloud-to-edge model compression with invisible data and models hosted on the server. B2KD faces challenges such as limited Internet exchange and edge-cloud disparity of data distributions. In this paper, we formalize a two-step workflow consisting of deprivatization and distillation, and theoretically provide a new optimization direction from logits to cell boundary different from direct logits alignment. With its guidance, we propose a new method Mapping-Emulation KD (MEKD) that distills a black-box cumbersome model into a lightweight one. Our method does not differentiate between treating soft or hard responses, and consists of: 1) deprivatization: emulating the inverse mapping of the teacher function with a generator, and 2) distillation: aligning low-dimensional logits of the teacher and student models by reducing the distance of high-dimensional image points. For different teacher-student pairs, our method yields inspiring distillation performance on various benchmarks, and outperforms the previous state-of-the-art approaches.
翻译:黑盒知识蒸馏(B2KD)是面向云到端模型压缩的公式化问题,其场景涉及服务器端不可见数据与模型。B2KD面临互联网交换受限以及云与端数据分布差异等挑战。本文形式化了一个由去私有化与蒸馏构成的两步工作流,并从理论上提出了一种不同于直接对齐对数、而着眼于决策边界的新优化方向。在该理论指导下,我们提出映射仿真知识蒸馏(MEKD)方法,将黑盒复杂模型蒸馏为轻量模型。该方法不区分软/硬响应处理,具体包括:1)去私有化阶段:通过生成器仿真教师函数的逆映射;2)蒸馏阶段:通过缩小高维图像点距离来实现教师与学生模型低维对数的对齐。在多种师生模型组合下,本方法在多个基准测试中展现出显著的蒸馏性能,并超越了此前最先进的方法。