Subword tokenization introduces a computational layer in language models where many distinct token sequences decode to the same surface form and preserve meaning, yet induce different internal computations. Despite this non-uniqueness, language models are typically trained using a single canonical longest-prefix tokenization. We formalize homotokens-alternative valid subword segmentations of the same lexical item-as a strictly meaning-preserving form of data augmentation. We introduce a lightweight training architecture that conditions canonical next-token prediction on sampled homotoken variants via an auxiliary causal encoder and block-causal cross-attention, without modifying the training objective or token interface. In data-constrained pretraining, homotoken augmentation consistently delays overfitting under repeated data exposure and improves generalization across diverse evaluation datasets. In multilingual fine-tuning, we find that the effectiveness of homotokens depends on tokenizer quality: gains are strongest when canonical tokens are highly compressed and diminish when the tokenizer already over-fragments the input. Overall, homotokens provide a simple and modular mechanism for inducing tokenization invariance in language models.


翻译:子词分词在语言模型中引入了一个计算层,其中许多不同的分词序列会解码为相同的表层形式并保持语义,但会引发不同的内部计算。尽管存在这种非唯一性,语言模型通常使用单一规范的最长前缀分词进行训练。我们将同形分词——同一词汇项的有效替代子词切分——形式化为一种严格保持语义的数据增强方法。我们提出一种轻量级训练架构,通过辅助因果编码器和块因果交叉注意力,将规范的下一个分词预测条件化于采样的同形分词变体,而无需修改训练目标或分词接口。在数据受限的预训练中,同形分词增强能持续延迟重复数据暴露下的过拟合现象,并在多样化评估数据集上提升泛化能力。在多语言微调场景中,我们发现同形分词的有效性取决于分词器质量:当规范分词具有高度压缩性时增益最强,而当分词器已对输入进行过度分割时增益减弱。总体而言,同形分词为在语言模型中引入分词不变性提供了一种简单且模块化的机制。

0
下载
关闭预览

相关内容

将一个汉字序列切分成一个一个单独的词
大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
零训练开放词汇语义分割综述
专知会员服务
11+阅读 · 2025年5月31日
什么是后训练?大语言模型训练后优化方法综述,87页pdf
预训练语言模型的应用综述
专知会员服务
36+阅读 · 2023年1月23日
EMNLP 2021 | 预训练跨语言模型中的大词表构建及使用
专知会员服务
22+阅读 · 2022年1月5日
绝对干货!NLP预训练模型:从transformer到albert
新智元
15+阅读 · 2019年11月10日
如何理解模型的过拟合与欠拟合,以及如何解决?
七月在线实验室
12+阅读 · 2019年4月23日
NLP预训练模型大集合!
机器之心
21+阅读 · 2018年12月28日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
深度学习 | 利用词嵌入对文本进行情感分析
沈浩老师
11+阅读 · 2017年10月19日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
《人工智能赋能的适应性多功能电磁战》
专知会员服务
9+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
7+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
9+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
14+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
10+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
9+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
14+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
17+阅读 · 9月22日
相关资讯
绝对干货!NLP预训练模型:从transformer到albert
新智元
15+阅读 · 2019年11月10日
如何理解模型的过拟合与欠拟合,以及如何解决?
七月在线实验室
12+阅读 · 2019年4月23日
NLP预训练模型大集合!
机器之心
21+阅读 · 2018年12月28日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
深度学习 | 利用词嵌入对文本进行情感分析
沈浩老师
11+阅读 · 2017年10月19日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员