Although recent neural models for coreference resolution have led to substantial improvements on benchmark datasets, transferring these models to new target domains containing out-of-vocabulary spans and requiring differing annotation schemes remains challenging. Typical approaches involve continued training on annotated target-domain data, but obtaining annotations is costly and time-consuming. We show that annotating mentions alone is nearly twice as fast as annotating full coreference chains. Accordingly, we propose a method for efficiently adapting coreference models, which includes a high-precision mention detection objective and requires annotating only mentions in the target domain. Extensive evaluation across three English coreference datasets: CoNLL-2012 (news/conversation), i2b2/VA (medical notes), and previously unstudied child welfare notes, reveals that our approach facilitates annotation-efficient transfer and results in a 7-14% improvement in average F1 without increasing annotator time.
翻译:尽管近期基于神经网络的共指消解模型在基准数据集上取得了显著改进,但将这些模型迁移到包含词汇表外片段且需要不同标注方案的新目标领域仍具挑战。典型方法涉及对标注的目标领域数据进行持续训练,但获取标注成本高昂且耗时。我们证明,仅标注提及的速度几乎是标注完整共指链的两倍。据此,我们提出一种高效适应共指模型的方法,该方法包含高精度提及检测目标,且仅需标注目标领域的提及。在三个英文共指数据集(CoNLL-2012(新闻/对话)、i2b2/VA(医疗记录)及此前未研究的儿童福利记录)上的广泛评估显示,我们的方法实现了标注高效的迁移,并在不增加标注者时间的情况下,平均F1值提升了7-14%。