We conducted a human subject study of named entity recognition on a noisy corpus of conversational music recommendation queries, with many irregular and novel named entities. We evaluated the human NER linguistic behaviour in these challenging conditions and compared it with the most common NER systems nowadays, fine-tuned transformers. Our goal was to learn about the task to guide the design of better evaluation methods and NER algorithms. The results showed that NER in our context was quite hard for both human and algorithms under a strict evaluation schema; humans had higher precision, while the model higher recall because of entity exposure especially during pre-training; and entity types had different error patterns (e.g. frequent typing errors for artists). The released corpus goes beyond predefined frames of interaction and can support future work in conversational music recommendation.
翻译:我们针对存在大量不规则和新颖命名实体的嘈杂会话式音乐推荐查询语料库,开展了一项人体受试者的命名实体识别研究。我们评估了人类在挑战性条件下的NER语言行为,并将其与当前最常用的NER系统——即经过微调的Transformer模型进行了比较。我们的目标是通过探究该任务来指导更优评估方法与NER算法的设计。结果表明,在严格评估框架下,我们的情境中NER对人类和算法都颇具难度;人类具有更高的精确率,而模型由于预训练阶段对实体的暴露(尤其是实体曝光),拥有更高的召回率;此外,不同实体类型呈现出不同的错误模式(例如,艺术家名称频繁出现的拼写错误)。该发布语料库超越了预定义的交互框架,可为未来会话式音乐推荐的相关研究提供支持。