In this paper, we present our system for the BioNNE English track, which aims to extract 8 types of biomedical nested named entities from biomedical text. We use a large language model (Mixtral 8x7B instruct) and ScispaCy NER model to identify entities in an article and build custom heuristics based on unified medical language system (UMLS) semantic types to categorize the entities. We discuss the results and limitations of our system and propose future improvements. Our system achieved an F1 score of 0.39 on the BioNNE validation set and 0.348 on the test set.
翻译:本文提出用于BioNNE英文赛道任务的系统,其目标是从生物医学文本中抽取8类生物医学嵌套命名实体。我们采用大语言模型(Mixtral 8x7B instruct)与ScispaCy NER模型识别文献中的实体,并基于统一医学语言系统(UMLS)语义类型构建定制化启发式规则以完成实体分类。文中讨论了该系统的实验结果与局限性,并提出了未来改进方向。本系统在BioNNE验证集上获得0.39的F1分数,在测试集上获得0.348的F1分数。