End-to-end relation extraction (E2ERE) is an important task in information extraction, more so for biomedicine as scientific literature continues to grow exponentially. E2ERE typically involves identifying entities (or named entity recognition (NER)) and associated relations, while most RE tasks simply assume that the entities are provided upfront and end up performing relation classification. E2ERE is inherently more difficult than RE alone given the potential snowball effect of errors from NER leading to more errors in RE. A complex dataset in biomedical E2ERE is the ChemProt dataset (BioCreative VI, 2017) that identifies relations between chemical compounds and genes/proteins in scientific literature. ChemProt is included in all recent biomedical natural language processing benchmarks including BLUE, BLURB, and BigBio. However, its treatment in these benchmarks and in other separate efforts is typically not end-to-end, with few exceptions. In this effort, we employ a span-based pipeline approach to produce a new state-of-the-art E2ERE performance on the ChemProt dataset, resulting in $> 4\%$ improvement in F1-score over the prior best effort. Our results indicate that a straightforward fine-grained tokenization scheme helps span-based approaches excel in E2ERE, especially with regards to handling complex named entities. Our error analysis also identifies a few key failure modes in E2ERE for ChemProt.


翻译:端到端关系抽取(E2ERE)是信息抽取领域的重要任务,在生物医学文献呈指数级增长的背景下尤为关键。E2ERE通常涉及实体识别(即命名实体识别(NER))及其关联关系检测,而大多数关系抽取任务则直接假设实体已预先给定,仅需完成关系分类。由于NER阶段的错误可能产生连锁反应导致关系抽取错误率上升,E2ERE本质上比单纯的关系抽取更具挑战性。生物医学E2ERE领域的典型复杂数据集是ChemProt(BioCreative VI, 2017),该数据集用于识别科学文献中化合物与基因/蛋白质之间的关联关系。该数据集已被纳入BLUE、BLURB和BigBio等最新生物医学自然语言处理基准测试,但除少数例外情况外,这些基准测试及其他独立研究中通常未采用端到端方法处理该数据集。本研究采用基于跨度的流水线方法,在ChemProt数据集上实现了新的最优端到端关系抽取性能,F1分数较先前最佳方法提升超过4%。实验结果表明,简单的细粒度分词方案有助于基于跨度的方法在端到端关系抽取中取得卓越表现,尤其在处理复杂命名实体时效果显著。通过错误分析,我们还识别出ChemProt数据集中端到端关系抽取的若干关键失效模式。

0
下载
关闭预览

相关内容

Nat. Biotechnol. | 用机器学习预测多肽质谱库
专知会员服务
18+阅读 · 2022年9月12日
基于几何结构预训练的蛋白质表征学习
专知会员服务
15+阅读 · 2022年8月21日
AlphaFold预测出2亿种蛋白质结构,打开整个蛋白质宇宙
专知会员服务
14+阅读 · 2022年8月1日
【知识图谱@EMNLP2020】Knowledge Graphs in NLP @ EMNLP 2020
专知会员服务
43+阅读 · 2020年11月22日
GNN 新基准!Long Range Graph Benchmark
图与推荐
0+阅读 · 2022年10月18日
VCIP 2022 Call for Demos
CCF多媒体专委会
1+阅读 · 2022年6月6日
【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理
深度学习自然语言处理
18+阅读 · 2020年5月22日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
笔记 | Sentiment Analysis
黑龙江大学自然语言处理实验室
10+阅读 · 2018年5月6日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
2+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
1+阅读 · 2008年12月31日
Arxiv
0+阅读 · 2023年5月22日
Arxiv
13+阅读 · 2017年12月5日
VIP会员
最新内容
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
2+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
《基于强化学习的自动化红队测试》
专知会员服务
4+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
11+阅读 · 7月22日
相关VIP内容
Nat. Biotechnol. | 用机器学习预测多肽质谱库
专知会员服务
18+阅读 · 2022年9月12日
基于几何结构预训练的蛋白质表征学习
专知会员服务
15+阅读 · 2022年8月21日
AlphaFold预测出2亿种蛋白质结构,打开整个蛋白质宇宙
专知会员服务
14+阅读 · 2022年8月1日
【知识图谱@EMNLP2020】Knowledge Graphs in NLP @ EMNLP 2020
专知会员服务
43+阅读 · 2020年11月22日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
2+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
1+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员