To advance Chinese financial natural language processing (NLP), we introduce BBT-FinT5, a new Chinese financial pre-training language model based on the T5 model. To support this effort, we have built BBT-FinCorpus, a large-scale financial corpus with approximately 300GB of raw text from four different sources. In general domain NLP, comprehensive benchmarks like GLUE and SuperGLUE have driven significant advancements in language model pre-training by enabling head-to-head comparisons among models. Drawing inspiration from these benchmarks, we propose BBT-CFLEB, a Chinese Financial Language understanding and generation Evaluation Benchmark, which includes six datasets covering both understanding and generation tasks. Our aim is to facilitate research in the development of NLP within the Chinese financial domain. Our model, corpus and benchmark are released at https://github.com/ssymmetry/BBT-FinCUGE-Applications. Our work belongs to the Big Bang Transformer (BBT), a large-scale pre-trained language model project.
翻译:为推进中文金融自然语言处理(NLP)研究,我们基于T5模型提出了BBT-FinT5,一种全新的中文金融预训练语言模型。为支撑此项工作,我们构建了BBT-FinCorpus,一个包含约300GB原始文本的大规模金融语料库,其数据来源涵盖四个不同渠道。在通用领域NLP中,GLUE和SuperGLUE等综合基准通过支持模型间的直接对比,显著推动了语言模型预训练的发展。受这些基准启发,我们提出了BBT-CFLEB(中文金融语言理解与生成评估基准),该基准包含六个数据集,覆盖理解与生成两类任务。我们的目标是促进中文金融领域NLP研究的发展。我们的模型、语料库及基准已发布于https://github.com/ssymmetry/BBT-FinCUGE-Applications。本工作属于大爆炸Transformer(BBT)大规模预训练语言模型项目。