Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late or degrade utility for legitimate users. We propose \textbf{Knowledge Trap}, a defense that redirects extraction attacks toward low-transferability knowledge through a \emph{Honeypot Knowledge Graph} (HKG) and breadcrumb-guided exploration. Instead of blocking queries or perturbing outputs, Knowledge Trap consumes the attacker's limited query budget on knowledge with negligible downstream utility while preserving benign-user performance. Experiments in medical and financial domains show that Knowledge Trap reduces surrogate Agreement by 6.2\% on average without degrading legitimate-user accuracy, outperforming existing defenses that impose measurable user impact. These results suggest that defending knowledge-space traversal is a practical direction for mitigating LLM extraction attacks.
翻译:作为商业API部署的大语言模型易受模型提取攻击,而现有防御措施要么行动迟缓,要么降低了合法用户的效用。我们提出**知识陷阱**(Knowledge Trap)防御机制,通过*蜜罐知识图谱*(Honeypot Knowledge Graph, HKG)和面包屑引导式探索,将提取攻击引向低迁移性知识。知识陷阱并非阻止查询或扰动输出,而是消耗攻击者在低下游效用知识上的有限查询预算,同时保障良性用户性能。在医疗和金融领域的实验表明,知识陷阱在不降低合法用户准确率的前提下,平均使替代模型一致性降低6.2%,优于对用户造成可测量影响的现有防御措施。这些结果表明,防御知识空间遍历是缓解LLM提取攻击的可行方向。