As LLMs become more pervasive across various users and scenarios, identifying potential issues when using these models becomes essential. Examples include bias, inconsistencies, and hallucination. Although auditing the LLM for these problems is desirable, it is far from being easy or solved. An effective method is to probe the LLM using different versions of the same question. This could expose inconsistencies in its knowledge or operation, indicating potential for bias or hallucination. However, to operationalize this auditing method at scale, we need an approach to create those probes reliably and automatically. In this paper we propose an automatic and scalable solution, where one uses a different LLM along with human-in-the-loop. This approach offers verifiability and transparency, while avoiding circular reliance on the same LLMs, and increasing scientific rigor and generalizability. Specifically, we present a novel methodology with two phases of verification using humans: standardized evaluation criteria to verify responses, and a structured prompt template to generate desired probes. Experiments on a set of questions from TruthfulQA dataset show that we can generate a reliable set of probes from one LLM that can be used to audit inconsistencies in a different LLM. The criteria for generating and applying auditing probes is generalizable to various LLMs regardless of the underlying structure or training mechanism.
翻译:随着大型语言模型在不同用户和场景中的广泛应用,识别使用这些模型时可能出现的问题变得至关重要,例如偏见、不一致性和幻觉。尽管对这些模型进行审计以发现上述问题具有现实需求,但该任务远非易行或已解决。一种有效方法是使用同一问题的不同变体对模型进行探测,这有助于揭示其在知识或运行中的不一致性,从而暗示潜在的偏见或幻觉风险。然而,要大规模实施这一审计方法,我们需要一种可靠且自动化的探测生成方案。本文提出了一种自动化、可扩展的解决方案,通过结合不同的大型语言模型与人在回路机制实现。该方法在避免对同一语言模型形成循环依赖的同时,提升了可验证性与透明度,并增强了科学严谨性和通用性。具体而言,我们提出了一种包含两阶段人工验证的新方法:采用标准化评估标准验证模型响应,以及利用结构化提示模板生成所需探测。基于TruthfulQA数据集中的一组问题进行的实验表明,我们能够从一个语言模型生成可靠的探测集,用于审计另一个语言模型的不一致性。该探测生成与应用准则具有通用性,可适配不同底层结构或训练机制的各种大型语言模型。