Positional encoding (PE) is widely viewed as necessary for transformers to process ordered sequences: without them, the next-token map appears permutation-invariant in its context tokens. This intuition underlies all prior universality results, which rely on positional information to prove that transformers with chain-of-thought can perform arbitrary computation, i.e., they are Turing complete. We revisit this belief in the regime most relevant to long-form reasoning, where generation proceeds through a finite sliding context window. Our opening perception is that the window mechanism itself (mildly) breaks the permutation symmetry. To distill and precisely capture the degree of this added expressiveness, we introduce an abstract autoregressive model, the HIST model, in which each update depends only on constant-size internal state and the token-count histogram within the current window. We prove that this HIST model is Turing complete by showing that the evolution of the window can reveal the token that has just left the window, which suffices to simulate Turing-complete Post machines. We then construct a sliding-window transformer over a constant-size token alphabet, without PE, and show that it can simulate the HIST model. Our result demonstrates that positional encodings are not indispensable for transformers to perform universal computation: The window sliding itself already breaks permutation symmetry and captures sufficient positional information.


翻译:位置编码(PE)被广泛视为Transformer处理有序序列的必要条件:若无位置编码,逐令牌映射在其上下文令牌中呈现置换不变性。这一直觉支撑了所有先前的通用性结论——这些结论依赖位置信息来证明具有思维链的Transformer能执行任意计算(即达到图灵完备性)。本文在与长程推理最相关的场景中重新审视这一信念:此时生成过程通过有限滑动上下文窗口进行。我们的初始观察发现,窗口机制本身(轻微地)打破了置换对称性。为了提炼并精确刻画这种新增表达能力的程度,我们引入了一种抽象自回归模型——HIST模型,其中每次更新仅依赖恒定大小的内部状态与当前窗口内的令牌计数直方图。我们通过证明窗口演化能揭示刚离开窗口的令牌(这足以模拟图灵完备的Post机),证得HIST模型具有图灵完备性。随后,我们在恒定大小令牌字母表上构建无PE的滑动窗口Transformer,并证明它能模拟HIST模型。我们的结果表明,Transformer实现通用计算并不必需位置编码:窗口滑动本身已能打破置换对称性并捕获充分的位序信息。

0
下载
关闭预览

相关内容

Transformer的无限之路:位置编码视角下的长度外推综述
专知会员服务
44+阅读 · 2024年1月17日
144页ppt!《Transformers》全面讲解,附视频
专知会员服务
119+阅读 · 2023年1月1日
专知会员服务
30+阅读 · 2021年7月30日
【ICML2021】具有线性复杂度的Transformer的相对位置编码
专知会员服务
25+阅读 · 2021年5月20日
Transformer文本分类代码
专知会员服务
118+阅读 · 2020年2月3日
【Tutorial】计算机视觉中的Transformer,98页ppt
专知
21+阅读 · 2021年10月25日
从头开始了解Transformer
AI科技评论
25+阅读 · 2019年8月28日
百闻不如一码!手把手教你用Python搭一个Transformer
大数据文摘
18+阅读 · 2019年4月22日
多图带你读懂 Transformers 的工作原理
AI研习社
10+阅读 · 2019年3月18日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 3月25日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
6+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
5+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
9+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
11+阅读 · 8月8日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员