We propose the application of Transformer-based language models for classifying entity legal forms from raw legal entity names. Specifically, we employ various BERT variants and compare their performance against multiple traditional baselines. Our evaluation encompasses a substantial subset of freely available Legal Entity Identifier (LEI) data, comprising over 1.1 million legal entities from 30 different legal jurisdictions. The ground truth labels for classification per jurisdiction are taken from the Entity Legal Form (ELF) code standard (ISO 20275). Our findings demonstrate that pre-trained BERT variants outperform traditional text classification approaches in terms of F1 score, while also performing comparably well in the Macro F1 Score. Moreover, the validity of our proposal is supported by the outcome of third-party expert reviews conducted in ten selected jurisdictions. This study highlights the significant potential of Transformer-based models in advancing data standardization and data integration. The presented approaches can greatly benefit financial institutions, corporations, governments and other organizations in assessing business relationships, understanding risk exposure, and promoting effective governance.
翻译:我们提出应用基于Transformer的语言模型对原始法人实体名称进行法律形式分类。具体而言,我们采用多种BERT变体,并将其与多种传统基线方法的性能进行对比。评估使用了大量公开可用的法律实体标识符(LEI)数据子集,涵盖来自30个不同法域的110多万个法律实体。各法域的分类真实标签采用实体法律形式(ELF)代码标准(ISO 20275)。研究结果表明,预训练的BERT变体在F1分数上优于传统文本分类方法,同时在宏平均F1分数上也表现相当。此外,在十个选定法域开展的第三方专家评审结果进一步验证了本方案的有效性。本研究凸显了基于Transformer模型在推进数据标准化与数据集成方面的巨大潜力,所提出的方法可有力帮助金融机构、企业、政府及其他组织评估商业关系、理解风险敞口并促进有效治理。