Analogy-making is central to human cognition, allowing us to adapt to novel situations -- an ability that current AI systems still lack. Most analogy datasets today focus on simple analogies (e.g., word analogies); datasets including complex types of analogies are typically manually curated and very small. We believe that this holds back progress in computational analogy. In this work, we design a data generation pipeline, ParallelPARC (Parallel Paragraph Creator) leveraging state-of-the-art Large Language Models (LLMs) to create complex, paragraph-based analogies, as well as distractors, both simple and challenging. We demonstrate our pipeline and create ProPara-Logy, a dataset of analogies between scientific processes. We publish a gold-set, validated by humans, and a silver-set, generated automatically. We test LLMs' and humans' analogy recognition in binary and multiple-choice settings, and found that humans outperform the best models (~13% gap) after a light supervision. We demonstrate that our silver-set is useful for training models. Lastly, we show challenging distractors confuse LLMs, but not humans. We hope our pipeline will encourage research in this emerging field.
翻译:类比推理是人类认知的核心能力,使我们能够适应新情境——而当前的人工智能系统仍缺乏这一能力。现今大多数类比数据集侧重于简单类比(如词汇类比),包含复杂类比类型的数据集通常依赖人工整理且规模极小。我们认为这种现状阻碍了计算类比领域的发展。本研究设计了名为ParallelPARC(并行段落生成器)的数据生成流程,该流程利用最先进的大语言模型(LLMs)构建复杂段落级类比,并生成简单与高难度两种干扰项。我们通过该流程创建了ProPara-Logy数据集——科学过程之间的类比数据集,并发布了经人工验证的金标集与自动生成的银标集。在二元分类和多项选择的类比识别测试中,经简单监督后人类表现优于最佳模型(约13%差距)。实验证明银标集对模型训练具有实用价值,同时显示高难度干扰项会混淆大语言模型,但不会影响人类判断。我们期望本数据流程能推动这一新兴领域的研究。