Transformer-based masked language models such as BERT, trained on general corpora, have shown impressive performance on downstream tasks. It has also been demonstrated that the downstream task performance of such models can be improved by pretraining larger models for longer on more data. In this work, we empirically evaluate the extent to which these results extend to tasks in science. We use 14 domain-specific transformer-based models (including ScholarBERT, a new 770M-parameter science-focused masked language model pretrained on up to 225B tokens) to evaluate the impact of training data, model size, pretraining and finetuning time on 12 downstream scientific tasks. Interestingly, we find that increasing model sizes, training data, or compute time does not always lead to significant improvements (i.e., >1% F1), if at all, in scientific information extraction tasks and offered possible explanations for the surprising performance differences.
翻译:基于Transformer的掩码语言模型(如BERT)在通用语料上训练后,在下游任务中展现出卓越性能。研究表明,通过扩大模型规模、延长训练时间或增加训练数据,可进一步提升此类模型的下游任务表现。本研究通过实证方法系统评估了这些结论在科学领域任务中的泛化程度。我们采用14个领域特定的Transformer模型(包括新提出的ScholarBERT——一个参数规模达7.7亿、基于2250亿词元训练的掩码语言模型),系统评估训练数据量、模型规模、预训练及微调时长对12项科学下游任务的影响。有趣的是,实验发现增加模型规模、训练数据量或计算资源在科学信息抽取任务中并不总能带来显著性能提升(即F1值增幅超过1%),并针对这一意外性能差异提出了可能的解释。