Medical decision-making processes can be enhanced by comprehensive biomedical knowledge bases, which require fusing knowledge graphs constructed from different sources via a uniform index system. The index system often organizes biomedical terms in a hierarchy to provide the aligned entities with fine-grained granularity. To address the challenge of scarce supervision in the biomedical knowledge fusion (BKF) task, researchers have proposed various unsupervised methods. However, these methods heavily rely on ad-hoc lexical and structural matching algorithms, which fail to capture the rich semantics conveyed by biomedical entities and terms. Recently, neural embedding models have proved effective in semantic-rich tasks, but they rely on sufficient labeled data to be adequately trained. To bridge the gap between the scarce-labeled BKF and neural embedding models, we propose HiPrompt, a supervision-efficient knowledge fusion framework that elicits the few-shot reasoning ability of large language models through hierarchy-oriented prompts. Empirical results on the collected KG-Hi-BKF benchmark datasets demonstrate the effectiveness of HiPrompt.
翻译:医疗决策过程可通过整合多方来源构建的知识图谱并借助统一索引体系来增强,该索引体系通常以层级结构组织生物医学术语,为对齐实体提供细粒度语义支持。针对生物医学知识融合任务中标注数据稀缺的挑战,研究者已提出多种无监督方法。然而,这些方法过度依赖特定领域的词汇与结构匹配算法,难以捕捉生物医学实体与术语蕴含的丰富语义。近年来,神经嵌入模型在语义密集型任务中展现出有效性,但其充分训练需要充足的标注数据。为弥合标注稀缺型生物医学知识融合与神经嵌入模型之间的鸿沟,我们提出HiPrompt——一种高效监督的知识融合框架,通过面向层次的大语言模型提示激发其小样本推理能力。在收集的KG-Hi-BKF基准数据集上的实验结果表明了HiPrompt的有效性。