Pre-trained Transformers are challenging human performances in many NLP tasks. The massive datasets used for pre-training seem to be the key to their success on existing tasks. In this paper, we explore how a range of pre-trained Natural Language Understanding models perform on definitely unseen sentences provided by classification tasks over a DarkNet corpus. Surprisingly, results show that syntactic and lexical neural networks perform on par with pre-trained Transformers even after fine-tuning. Only after what we call extreme domain adaptation, that is, retraining with the masked language model task on all the novel corpus, pre-trained Transformers reach their standard high results. This suggests that huge pre-training corpora may give Transformers unexpected help since they are exposed to many of the possible sentences.
翻译:预训练Transformer在许多自然语言处理任务中挑战人类表现。用于预训练的大规模数据集似乎是在现有任务上成功的关键。本文探讨了一系列预训练自然语言理解模型在暗网语料库分类任务提供的明确未见句子上的表现。令人惊讶的是,结果显示即使经过微调,句法和词汇神经网络的表现也与预训练Transformer不相上下。只有在我们称之为极端领域适应——即使用掩码语言模型任务在所有新语料库上重新训练之后——预训练Transformer才能达到其标准高表现。这表明巨大的预训练语料库可能给Transformer带来意想不到的帮助,因为它们暴露在许多可能的句子中。