Analogical reasoning is a fundamental cognitive ability of humans. However, current language models (LMs) still struggle to achieve human-like performance in analogical reasoning tasks due to a lack of resources for model training. In this work, we address this gap by proposing ANALOGYKB, a million-scale analogy knowledge base (KB) derived from existing knowledge graphs (KGs). ANALOGYKB identifies two types of analogies from the KGs: 1) analogies of the same relations, which can be directly extracted from the KGs, and 2) analogies of analogous relations, which are identified with a selection and filtering pipeline enabled by large LMs (InstructGPT), followed by minor human efforts for data quality control. Evaluations on a series of datasets of two analogical reasoning tasks (analogy recognition and generation) demonstrate that ANALOGYKB successfully enables LMs to achieve much better results than previous state-of-the-art methods.
翻译:类比推理是人类的基本认知能力。然而,由于缺乏模型训练资源,当前语言模型在类比推理任务中仍难以达到人类水平的性能。本研究通过提出ANALOGYKB——一个从现有知识图谱衍生出的百万级类比知识库——来填补这一空白。ANALOGYKB从知识图谱中识别两类类比:1)相同关系的类比,可直接从知识图谱中提取;2)类比关系的类比,通过结合大型语言模型(InstructGPT)的选择与过滤流程,辅以少量人工数据质量管控来识别。在两个类比推理任务(类比识别与生成)的系列数据集上的评估表明,ANALOGYKB成功使语言模型取得了远超先前最先进方法的效果。