Large Language Models (LLMs) exploit fine-tuning as a technique to adapt to diverse goals, thanks to task-specific training data. Task specificity should go hand in hand with domain orientation, that is, the specialization of an LLM to accurately address the tasks of a given realm of interest. However, models are usually fine-tuned over publicly available data or, at most, over ground data from databases, ignoring business-level definitions and domain experience. On the other hand, Enterprise Knowledge Graphs (EKGs) are able to capture and augment such domain knowledge via ontological reasoning. With the goal of combining LLM flexibility with the domain orientation of EKGs, we propose a novel neurosymbolic architecture that leverages the power of ontological reasoning to build task- and domain-specific corpora for LLM fine-tuning.
翻译:大型语言模型(LLMs)利用微调技术,通过任务特定的训练数据来适应多样化目标。任务特异性应与领域导向相结合,即语言模型针对特定兴趣领域任务进行精确处理的专业能力。然而,现有模型通常基于公开数据或最多基于数据库中的原始数据进行微调,忽略了业务层面的定义与领域经验。另一方面,企业知识图谱(EKGs)能够通过本体推理捕获并增强此类领域知识。为融合LLMs的灵活性与EKGs的领域导向特性,我们提出一种新颖的神经符号架构,该架构利用本体推理的力量构建面向任务和领域的专用语料库,用于LLMs的微调。