Despite the strong performance of large language models (LLMs) across a wide range of tasks, they still have reliability issues. Previous studies indicate that strong LLMs like GPT-4-turbo excel in evaluating the reliability of responses from LLMs, but face efficiency and local deployment issues. Thus, to enable weak LLMs to effectively assess the reliability of LLM responses, we propose a novel cross-query-comparison-based method called $\textit{Meta Ranking}$ (MR). Unlike previous few-shot methods that solely based on in-context learning capabilities in LLMs, MR assesses reliability by pairwisely ranking the target query-response pair with multiple reference query-response pairs. We found that MR is highly effective in error detection for LLM responses, where weak LLMs, such as Phi-2, could surpass strong baselines like GPT-3.5-turbo, requiring only five reference samples and significantly improving efficiency. We further demonstrate that MR can enhance strong LLMs' performance in two practical applications: model cascading and instruction tuning. In model cascading, we combine open- and closed-source LLMs to achieve performance comparable to GPT-4-turbo with lower costs. In instruction tuning, we use MR for iterative training data filtering, significantly reducing data processing time and enabling LLaMA-7B and Phi-2 to surpass Alpaca-13B with fewer training tokens. These results underscore the high potential of MR in both efficiency and effectiveness.
翻译:尽管大型语言模型(LLM)在广泛任务中表现出色,其可靠性问题依然存在。先前研究表明,如GPT-4-turbo等强LLM在评估LLM响应可靠性方面表现优异,但面临效率与本地部署的挑战。为此,我们提出一种基于跨查询比较的新方法——$\textit{元排序}$(MR),使弱LLM能有效评估LLM响应的可靠性。与以往仅依赖LLM上下文学习能力的少样本方法不同,MR通过将目标查询-响应对与多个参考查询-响应对进行成对排序来评估可靠性。我们发现MR在LLM响应错误检测方面极为有效:仅需五个参考样本,Phi-2等弱LLM即可超越GPT-3.5-turbo等强基线模型,并显著提升效率。我们进一步证明MR能在两个实际应用中增强强LLM的性能:模型级联与指令微调。在模型级联中,我们结合开源与闭源LLM,以更低成本实现与GPT-4-turbo相当的性能;在指令微调中,我们利用MR进行迭代训练数据过滤,大幅减少数据处理时间,使LLaMA-7B和Phi-2以更少的训练词元超越Alpaca-13B。这些结果充分证明了MR在效率与效能方面的巨大潜力。