Pretrained Transformers can perform in-context learning (ICL) from a few demonstrations, but this ability can fail sharply when the test distribution differs from pretraining, a common deployment setting. We study attention temperature as a simple inference-time control for improving ICL robustness under such shifts. In a high-dimensional linear-regression framework, we analyze a Transformer with "approximate softmax" attention, which preserves softmax's normalization and temperature-dependent selectivity while remaining tractable. We derive a closed-form expression for the ICL generalization error under distribution shift, and show that it is minimized by an explicit optimal attention temperature. This characterization yields interpretable guidance by linking the best temperature to moments of the pre-softmax attention scores, and predicts when temperature adjustment can recover near Bayes-optimal performance. We validate the theory with extensive simulations, and further demonstrate gains on pretrained LLMs (GPT-2 and Llama2-7B) on question-answering benchmarks under distribution shift induced by noisy in-context demonstrations. Overall, attention temperature emerges as a principled, lightweight knob for improving the robustness of ICL in pretrained Transformers.


翻译:预训练Transformer能从少量示例中进行上下文学习(ICL),但在测试分布与预训练分布不同这一常见部署场景下,这种能力可能急剧失效。本文研究注意力温度作为一种简单的推理时调控手段,以增强此类偏移下ICL的鲁棒性。在高维线性回归框架下,我们分析了一种采用“近似softmax”注意力的Transformer,该机制在保持可解性的同时,保留了softmax的归一化特性及温度依赖的选择性。我们推导出分布偏移下ICL泛化误差的闭合表达式,并证明其通过显式最优注意力温度达到最小化。这一表征通过将最优温度与预softmax注意力得分的矩相关联,提供了可解释的指导,并预测了何时温度调整能恢复接近贝叶斯最优的性能。我们通过大量仿真验证了该理论,并在噪声上下文示例引发的分布偏移下,基于预训练大语言模型(GPT-2和Llama2-7B)的问答基准测试中进一步展示了其性能提升。总体而言,注意力温度被视为一种原理简洁、轻量级的调控手段,可用于提升预训练Transformer中ICL的鲁棒性。

0
下载
关闭预览

相关内容

【NeurIPS2024】注意力迁移对视觉Transformer的惊人有效性研究
深度学习的下一步:Transformer和注意力机制
云头条
56+阅读 · 2019年9月14日
干货!自然语言处理中的自注意力机制!
全球人工智能
11+阅读 · 2018年3月27日
迁移学习在深度学习中的应用
专知
24+阅读 · 2017年12月24日
深度学习中的注意力机制
CSDN大数据
24+阅读 · 2017年11月2日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Arxiv
0+阅读 · 6月11日
Arxiv
0+阅读 · 6月9日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
8+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
14+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
相关VIP内容
【NeurIPS2024】注意力迁移对视觉Transformer的惊人有效性研究
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员