In recent years there has been a growing demand from financial agents, especially from particular and institutional investors, for companies to report on climate-related financial risks. A vast amount of information, in text format, can be expected to be disclosed in the short term by firms in order to identify these types of risks in their financial and non financial reports, particularly in response to the growing regulation that is being passed on the matter. To this end, this paper applies state-of-the-art NLP techniques to achieve the detection of climate change in text corpora. We use transfer learning to fine-tune two transformer models, BERT and ClimateBert -a recently published DistillRoBERTa-based model that has been specifically tailored for climate text classification-. These two algorithms are based on the transformer architecture which enables learning the contextual relationships between words in a text. We carry out the fine-tuning process of both models on the novel Clima-Text database, consisting of data collected from Wikipedia, 10K Files Reports and web-based claims. Our text classification model obtained from the ClimateBert fine-tuning process on ClimaText, outperforms the models created with BERT and the current state-of-the-art transformer in this particular problem. Our study is the first one to implement on the ClimaText database the recently published ClimateBert algorithm. Based on our results, it can be said that ClimateBert fine-tuned on ClimaText is an outstanding tool within the NLP pre-trained transformer models that may and should be used by investors, institutional agents and companies themselves to monitor the disclosure of climate risk in financial reports. In addition, our transfer learning methodology is cheap in computational terms, thus allowing any organization to perform it.
翻译:近年来,金融从业者(尤其是个人和机构投资者)日益要求企业报告气候相关金融风险。为应对日益严格的监管要求,企业有望在短期内披露大量文本格式信息,以在其财务和非财务报告中识别此类风险。为此,本文采用最先进的自然语言处理技术实现文本语料中气候变化信息的检测。我们运用迁移学习微调两种Transformer模型:BERT与ClimateBert——一种近期发布的基于DistillRoBERTa、专为气候文本分类定制的模型。这两种算法均基于Transformer架构,能够学习文本中单词的上下文关系。我们在新型Clima-Text数据库(包含来自维基百科、10K文件报告及网络声明的数据)上对两种模型进行微调。通过ClimateBert在ClimaText上的微调过程获得的文本分类模型,在此特定问题上优于基于BERT及当前最先进Transformer模型构建的结果。本研究率先在ClimaText数据库上应用近期发布的ClimateBert算法。基于实验结果,经ClimaText微调的ClimateBert堪称自然语言处理预训练Transformer模型中的杰出工具,可供投资者、机构主体及企业自身用于监测财务报告中的气候风险披露。此外,我们的迁移学习方法计算成本低廉,便于各类组织部署实施。