The disconnect between tokenizer creation and model training in language models has been known to allow for certain inputs, such as the infamous SolidGoldMagikarp token, to induce unwanted behaviour. Although such `glitch tokens' that are present in the tokenizer vocabulary, but are nearly or fully absent in training, have been observed across a variety of different models, a consistent way of identifying them has been missing. We present a comprehensive analysis of Large Language Model (LLM) tokenizers, specifically targeting this issue of detecting untrained and under-trained tokens. Through a combination of tokenizer analysis, model weight-based indicators, and prompting techniques, we develop effective methods for automatically detecting these problematic tokens. Our findings demonstrate the prevalence of such tokens across various models and provide insights into improving the efficiency and safety of language models.
翻译:语言模型中分词器创建与模型训练之间的脱节已知会导致某些输入(例如臭名昭著的“SolidGoldMagikarp”词元)引发异常行为。尽管这种存在于分词器词汇表中但在训练中几乎或完全缺失的“故障词元”已在多种不同模型中被观察到,但一直缺乏一致的识别方法。我们针对大语言模型(LLM)分词器开展了全面分析,专门聚焦于检测未训练和训练不足词元这一问题。通过结合分词器分析、基于模型权重的指标以及提示技术,我们开发出自动检测这些有问题词元的有效方法。我们的研究结果证明了此类词元在各种模型中的普遍性,并为提高语言模型的效率与安全性提供了见解。