While Large Language Models (LLMs) have achieved remarkable performance in many tasks, much about their inner workings remains unclear. In this study, we present novel experimental insights into the resilience of LLMs, particularly GPT-4, when subjected to extensive character-level permutations. To investigate this, we first propose the Scrambled Bench, a suite designed to measure the capacity of LLMs to handle scrambled input, in terms of both recovering scrambled sentences and answering questions given scrambled context. The experimental results indicate that most powerful LLMs demonstrate the capability akin to typoglycemia, a phenomenon where humans can understand the meaning of words even when the letters within those words are scrambled, as long as the first and last letters remain in place. More surprisingly, we found that only GPT-4 nearly flawlessly processes inputs with unnatural errors, even under the extreme condition, a task that poses significant challenges for other LLMs and often even for humans. Specifically, GPT-4 can almost perfectly reconstruct the original sentences from scrambled ones, decreasing the edit distance by 95%, even when all letters within each word are entirely scrambled. It is counter-intuitive that LLMs can exhibit such resilience despite severe disruption to input tokenization caused by scrambled text.
翻译:尽管大语言模型(LLMs)在许多任务中取得了卓越表现,但其内部运作机制仍不明确。本研究针对LLMs(尤其是GPT-4)在字符级全排列干扰下的鲁棒性提出了新颖的实验见解。为此,我们首先提出Scrambled Bench基准测试套件,该套件从恢复乱序句子和基于乱序上下文回答问题两个维度,系统评估LLMs处理乱序输入的能力。实验结果表明,多数强大LLMs展现出类似"典型词盲症"(typoglycemia)的能力——即人类即使单词内部字母顺序错乱,只要首尾字母位置不变仍可理解词义的现象。更令人惊讶的是,我们发现仅GPT-4能在极端条件下近乎完美地处理带有非自然错误的输入,这项任务对其他LLMs甚至人类都构成重大挑战。具体而言,即便每个单词内所有字母完全乱序,GPT-4仍可将乱序句子重建为原始句子,编辑距离降低95%。尽管乱序文本严重破坏了输入分词过程,LLMs却展现出如此韧性,这一现象与直觉相悖。