Proprietary large language models (LLMs) embody substantial economic value and are generally exposed only as black-box APIs, yet adversaries can still exploit their outputs to extract knowledge via distillation. Existing defenses focus exclusively on text-based distillation, leaving the important logit-based distillation largely unexplored. In this work, we analyze this problem and present an effective solution from an information-theoretic perspective. We characterize distillation-relevant information in teacher outputs using the conditional mutual information (CMI) between teacher logits and input queries conditioned on ground-truth labels. This quantity captures contextual information beneficial for model extraction, motivating us to defend distillation via CMI minimization. Guided by our theoretical analysis, we propose learning a transformation matrix that purifies the original outputs to enhance distillation resistance. We further derive a CMI-inspired anti-distillation objective to optimize this transformation, which effectively removes distillation-relevant information while preserving output utility. Extensive experiments across multiple LLMs and strong distillation algorithms demonstrate that the proposed method significantly degrades distillation performance while preserving task accuracy, effectively protecting models' intellectual property.


翻译:专有大语言模型(LLMs)蕴含巨大的经济价值,通常仅以黑盒API形式开放,然而攻击者仍可利用其输出通过蒸馏方法提取知识。现有防御机制仅专注于基于文本的蒸馏,而重要的基于逻辑值的蒸馏方法在很大程度上尚未得到充分探索。本研究从信息论角度分析该问题并提出一种有效的解决方案。我们利用教师模型逻辑值与输入查询在给定真实标签条件下的条件互信息(CMI)来刻画教师输出中与蒸馏相关的信息。该度量捕捉了有利于模型提取的上下文信息,从而启发我们通过CMI最小化来防御蒸馏。在理论分析的指导下,我们提出学习一个变换矩阵来纯化原始输出以增强抗蒸馏能力。进一步推导出受CMI启发的抗蒸馏目标函数来优化该变换,该方法能有效去除与蒸馏相关的信息同时保持输出效用。在多种LLMs和强蒸馏算法上的大量实验表明,所提方法在保持任务精度的同时能显著降低蒸馏性能,有效保护模型的知识产权。

0
下载
关闭预览

相关内容

面向统计学家的大型语言模型概述
专知会员服务
32+阅读 · 2025年3月16日
迈向大语言模型偏好学习的统一视角综述
专知会员服务
24+阅读 · 2024年9月7日
大型语言模型的知识蒸馏综述:方法、评估与应用
专知会员服务
81+阅读 · 2024年7月4日
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
模型压缩 | 知识蒸馏经典解读
AINLP
11+阅读 · 2020年5月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
7+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
0+阅读 · 16分钟前
《战略战术化:一项综合性述评》
专知会员服务
0+阅读 · 20分钟前
美陆军-工业界协同推进反无人机系统技术发展
专知会员服务
1+阅读 · 42分钟前
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
0+阅读 · 48分钟前
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
7+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员