Large Language Models (LLMs) have shown impressive abilities in data annotation, opening the way for new approaches to solve classic NLP problems. In this paper, we show how to use LLMs to create NuNER, a compact language representation model specialized in the Named Entity Recognition (NER) task. NuNER can be fine-tuned to solve downstream NER problems in a data-efficient way, outperforming similar-sized foundation models in the few-shot regime and competing with much larger LLMs. We find that the size and entity-type diversity of the pre-training dataset are key to achieving good performance. We view NuNER as a member of the broader family of task-specific foundation models, recently unlocked by LLMs.
翻译:摘要:大语言模型(LLMs)在数据标注方面展现出卓越能力,为经典自然语言处理问题的解决开辟了新路径。本文展示了如何利用LLMs构建NuNER——一种专攻命名实体识别(NER)任务的紧凑型语言表征模型。NuNER可通过微调在数据高效条件下解决下游NER问题,其在少样本场景中性能优于同等规模的基座模型,并可媲美参数规模更大的LLMs。研究发现,预训练数据集规模与实体类型多样性是达成优异性能的关键因素。我们将NuNER视为由LLMs最新驱动的任务专用基座模型大家族中的一员。