In this work, we conduct a detailed analysis on the performance of legal-oriented pre-trained language models (PLMs). We examine the interplay between their original objective, acquired knowledge, and legal language understanding capacities which we define as the upstream, probing, and downstream performance, respectively. We consider not only the models' size but also the pre-training corpora used as important dimensions in our study. To this end, we release a multinational English legal corpus (LeXFiles) and a legal knowledge probing benchmark (LegalLAMA) to facilitate training and detailed analysis of legal-oriented PLMs. We release two new legal PLMs trained on LeXFiles and evaluate them alongside others on LegalLAMA and LexGLUE. We find that probing performance strongly correlates with upstream performance in related legal topics. On the other hand, downstream performance is mainly driven by the model's size and prior legal knowledge which can be estimated by upstream and probing performance. Based on these findings, we can conclude that both dimensions are important for those seeking the development of domain-specific PLMs.
翻译:在本研究中,我们对面向法律领域的预训练语言模型(PLMs)的性能进行了详细分析。我们考察了这些模型的原始目标、获取的知识以及法律语言理解能力之间的相互作用,并将其分别定义为上游性能、探针性能和下游性能。我们在研究中不仅考虑了模型的规模,还将预训练语料库视为重要维度。为此,我们发布了一个多国英语法律语料库(LeXFiles)和一个法律知识探针基准(LegalLAMA),以促进面向法律领域的PLMs的训练和详细分析。我们发布了两个基于LeXFiles训练的新法律PLMs,并与其它模型在LegalLAMA和LexGLUE上进行了评估。研究发现,探针性能与相关法律主题的上游性能高度相关。另一方面,下游性能主要由模型规模和先验法律知识驱动,这些可通过上游和探针性能进行估计。基于这些发现,我们可以得出结论:对于寻求开发特定领域PLMs的研究者而言,这两个维度都至关重要。