We present a new pre-trained language model (PLM) for modern Hebrew, termed AlephBERTGimmel, which employs a much larger vocabulary (128K items) than standard Hebrew PLMs before. We perform a contrastive analysis of this model against all previous Hebrew PLMs (mBERT, heBERT, AlephBERT) and assess the effects of larger vocabularies on task performance. Our experiments show that larger vocabularies lead to fewer splits, and that reducing splits is better for model performance, across different tasks. All in all this new model achieves new SOTA on all available Hebrew benchmarks, including Morphological Segmentation, POS Tagging, Full Morphological Analysis, NER, and Sentiment Analysis. Subsequently we advocate for PLMs that are larger not only in terms of number of layers or training data, but also in terms of their vocabulary. We release the new model publicly for unrestricted use.
翻译:本文提出一个面向现代希伯来语的新型预训练语言模型(PLM),命名为AlephBERTGimmel,该模型采用比标准希伯来语PLM大得多的词表(128K个词项)。我们对该模型与所有此前希伯来语PLM(mBERT、heBERT、AlephBERT)进行对比分析,系统评估词表规模对下游任务性能的影响。实验表明,更大的词表能减少词汇切分数量,且减少切分数量对各类任务的模型性能均有益处。最终,该新模型在所有可用希伯来语基准测试中创下最新最优结果(SOTA),涵盖形态学分割、词性标注、完整形态分析、命名实体识别及情感分析。基于此,我们主张PLM的扩展不应仅局限于增加层数或训练数据量,还应扩大词表规模。我们公开发布该新模型以供无限制使用。