Geographical location is a crucial element of humanitarian response, outlining vulnerable populations, ongoing events, and available resources. Latest developments in Natural Language Processing may help in extracting vital information from the deluge of reports and documents produced by the humanitarian sector. However, the performance and biases of existing state-of-the-art information extraction tools are unknown. In this work, we develop annotated resources to fine-tune the popular Named Entity Recognition (NER) tools Spacy and roBERTa to perform geotagging of humanitarian texts. We then propose a geocoding method FeatureRank which links the candidate locations to the GeoNames database. We find that not only does the humanitarian-domain data improves the performance of the classifiers (up to F1 = 0.92), but it also alleviates some of the bias of the existing tools, which erroneously favor locations in the Western countries. Thus, we conclude that more resources from non-Western documents are necessary to ensure that off-the-shelf NER systems are suitable for the deployment in the humanitarian sector.
翻译:地理位置是人道主义响应的关键要素,用于标识脆弱人群、当前事件及可用资源。自然语言处理的最新进展有助于从人道主义领域大量报告和文档中提取关键信息。然而,现有最先进信息提取工具的性能和偏差尚不明确。本研究开发了标注资源,对主流命名实体识别(NER)工具Spacy和roBERTa进行微调,以执行人道主义文本的地理标注。随后提出地理编码方法FeatureRank,将候选地点与GeoNames数据库关联。研究发现:基于人道主义领域数据不仅提升了分类器性能(F1值最高达0.92),还缓解了现有工具偏向西方国家的错误偏差。因此,我们得出结论:为确保现成NER系统适用于人道主义领域部署,必须补充更多非西方文档资源。