Large Language Models (LLMs) have shown their success in language understanding and reasoning on general topics. However, their capability to inference based on user-specified structured data and knowledge in corpus-rare concepts like causal decision-making is still limited. In this work, we explore the possibility of fine-tuning an open-sourced LLM into LLM4Causal, which can identify the causal task, execute a corresponding function, and interpret its numerical results based on users' queries and the provided dataset. Meanwhile, we propose a data generation process for more controllable GPT prompting and present two instruction-tuning datasets: (1) Causal-Retrieval-Bench for causal problem identification and input parameter extraction for causal function calling and (2) Causal-Interpret-Bench for in-context causal interpretation. With three case studies, we showed that LLM4Causal can deliver end-to-end solutions for causal problems and provide easy-to-understand answers. Numerical studies also reveal that it has a remarkable ability to identify the correct causal task given a query.
翻译:大语言模型在通用主题的语言理解与推理方面已展现出成功,但其基于用户指定的结构化数据和像因果决策这类语料稀有的概念进行推理的能力仍然有限。在这项工作中,我们探索了将开源大语言模型微调为LLM4Causal的可能性,该模型能够识别因果任务、执行相应函数,并根据用户查询和提供的数据集解释其数值结果。同时,我们提出了一种用于更可控的GPT提示生成的数据生成过程,并提供了两个指令微调数据集:(1)用于因果问题识别和因果函数调用输入参数提取的Causal-Retrieval-Bench,(2)用于上下文因果解释的Causal-Interpret-Bench。通过三个案例研究,我们表明LLM4Causal能为因果问题提供端到端解决方案,并给出易理解的答案。数值研究还揭示,它具备根据查询识别正确因果任务的显著能力。