Large Language Models (LLMs), despite their impressive capabilities across domains, have been shown to be vulnerable to backdoor attacks. Prior backdoor strategies predominantly operate at the token level, where an injected trigger causes the model to generate a specific target word, choice, or class (depending on the task). Recent advances, however, exploit the long-form reasoning tendencies of modern LLMs to conduct reasoning-level backdoors: once triggered, the victim model inserts one or more malicious reasoning steps into its chain-of-thought (CoT). These attacks are substantially harder to detect, as the backdoored answer remains plausible and consistent with the poisoned reasoning trajectory. Yet, defenses tailored to this type of backdoor remain largely unexplored. To bridge this gap, we propose Critical-CoT, a novel defense mechanism that conducts a two-stage fine-tuning (FT) process on LLMs to develop critical thinking behaviors, enabling them to automatically identify potential backdoors and refuse to generate malicious reasoning steps. Extensive experiments across multiple LLMs and datasets demonstrate that Critical-CoT provides strong robustness against both in-context learning-based and FT-based backdoor attacks. Notably, Critical-CoT exhibits strong cross-domain and cross-task generalization. Our code is available at hthttps://github.com/tuanvu171/Critical-CoT.


翻译:尽管大型语言模型(LLMs)在各领域展现出卓越能力,但其已被证明易受后门攻击。以往的后门策略主要在词元级别运作,通过注入触发器使模型生成特定目标词、选项或类别(取决于任务)。然而,最新进展利用现代LLMs的长程推理倾向实现了推理级后门攻击:一旦触发,受害模型会向思维链(CoT)中插入一个或多个恶意推理步骤。由于后门答案仍保持合理性且与受污染的推理轨迹一致,此类攻击极难检测。然而,针对此类后门的防御机制仍鲜有探索。为弥补这一空白,我们提出Critical-CoT——一种新颖的防御机制,通过对LLMs进行两阶段微调(FT)培养其批判性思维行为,使其能够自动识别潜在后门并拒绝生成恶意推理步骤。在多种LLMs和数据集上的大量实验表明,Critical-CoT对基于上下文学习和基于FT的后门攻击均具有强鲁棒性。值得注意的是,Critical-CoT展现出优异的跨领域与跨任务泛化能力。我们的代码已开源至https://github.com/tuanvu171/Critical-CoT。

0
下载
关闭预览

相关内容

什么是后训练?大语言模型训练后优化方法综述,87页pdf
大型语言模型网络安全综述
专知会员服务
68+阅读 · 2024年5月12日
大型语言模型高效推理综述
专知会员服务
65+阅读 · 2024年4月23日
通信网络中大型语言模型的后门攻击的综述
专知会员服务
30+阅读 · 2023年9月5日
模型攻击:鲁棒性联邦学习研究的最新进展
机器之心
35+阅读 · 2020年6月3日
TheFatRat 一款简易后门工具
黑白之道
36+阅读 · 2019年10月23日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
1+阅读 · 今天14:52
《各指挥层级均亟需红队能力》报告
专知会员服务
2+阅读 · 今天14:46
《远程传感对战机架次生成影响的仿真研究》80页
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
7+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
9+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
5+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关VIP内容
相关基金
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员