While WangchanBERTa has become the de facto standard in transformer-based Thai language modeling, it still has shortcomings in regard to the understanding of foreign words, most notably English words, which are often borrowed without orthographic assimilation into Thai in many contexts. We identify the lack of foreign vocabulary in WangchanBERTa's tokenizer as the main source of these shortcomings. We then expand WangchanBERTa's vocabulary via vocabulary transfer from XLM-R's pretrained tokenizer and pretrain a new model using the expanded tokenizer, starting from WangchanBERTa's checkpoint, on a new dataset that is larger than the one used to train WangchanBERTa. Our results show that our new pretrained model, PhayaThaiBERT, outperforms WangchanBERTa in many downstream tasks and datasets.
翻译:尽管WangchanBERTa已成为基于Transformer的泰语语言建模事实标准,但其在处理外来词(尤其是英语词汇)时仍存在不足——这类词汇在许多语境中未经拼写同化直接借入泰语。我们发现WangchanBERTa分词器缺乏外来词汇是其缺陷的主要原因。为此,我们通过XLM-R预训练分词器的词汇迁移扩展WangchanBERTa的词典,并以WangchanBERTa检查点为初始权重,在规模更大的新数据集上利用扩展后的分词器预训练新模型。实验结果表明,我们的新预训练模型PhayaThaiBERT在多项下游任务与数据集上的表现均优于WangchanBERTa。