Large Language Models (LLMs) are extensively used today across various sectors, including academia, research, business, and finance, for tasks such as text generation, summarization, and translation. Despite their widespread adoption, these models often produce incorrect and misleading information, exhibiting a tendency to hallucinate. This behavior can be attributed to several factors, with consistency and reasoning capabilities being significant contributors. LLMs frequently lack the ability to generate explanations and engage in coherent reasoning, leading to inaccurate responses. Moreover, they exhibit inconsistencies in their outputs. This paper aims to evaluate and compare the consistency and reasoning capabilities of both public and proprietary LLMs. The experiments utilize the Boolq dataset as the ground truth, comprising questions, answers, and corresponding explanations. Queries from the dataset are presented as prompts to the LLMs, and the generated responses are evaluated against the ground truth answers. Additionally, explanations are generated to assess the models' reasoning abilities. Consistency is evaluated by repeatedly presenting the same query to the models and observing for variations in their responses. For measuring reasoning capabilities, the generated explanations are compared to the ground truth explanations using metrics such as BERT, BLEU, and F-1 scores. The findings reveal that proprietary models generally outperform public models in terms of both consistency and reasoning capabilities. However, even when presented with basic general knowledge questions, none of the models achieved a score of 90\% in both consistency and reasoning. This study underscores the direct correlation between consistency and reasoning abilities in LLMs and highlights the inherent reasoning challenges present in current language models.
翻译:大型语言模型(LLMs)如今广泛应用于学术、研究、商业和金融等多个领域,用于文本生成、摘要和翻译等任务。尽管它们被广泛采用,但这些模型经常产生错误和误导性信息,表现出产生幻觉的倾向。这种行为可归因于多个因素,其中一致性和推理能力是重要的贡献因素。LLMs 往往缺乏生成解释和进行连贯推理的能力,从而导致不准确的响应。此外,它们的输出存在不一致性。本文旨在评估和比较公共和专有 LLMs 的一致性和推理能力。实验使用 Boolq 数据集作为基准事实,该数据集包含问题、答案及相应的解释。将数据集中的查询作为提示输入给 LLMs,并将生成的响应与基准事实答案进行对比。此外,通过生成解释来评估模型的推理能力。一致性通过向模型重复呈现相同查询并观察其响应变化来评估。对于推理能力的测量,使用 BERT、BLEU 和 F-1 分数等指标将生成的解释与基准事实解释进行比较。研究结果表明,专有模型在一致性和推理能力方面通常优于公共模型。然而,即便是面对基础常识问题,所有模型在一致性和推理能力上均未达到 90% 的得分。这项研究强调了 LLMs 中一致性与推理能力之间的直接相关性,并揭示了当前语言模型中固有的推理挑战。