Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them in real-world applications to ensure they produce reliable performance. Despite the well-established importance of evaluating LLMs in the community, the complexity of the evaluation process has led to varied evaluation setups, causing inconsistencies in findings and interpretations. To address this, we systematically review the primary challenges and limitations causing these inconsistencies and unreliable evaluations in various steps of LLM evaluation. Based on our critical review, we present our perspectives and recommendations to ensure LLM evaluations are reproducible, reliable, and robust.
翻译:近年来,大规模语言模型(LLMs)因其在不同领域执行多样化任务的卓越能力而受到广泛关注。然而,在将这些模型部署于实际应用之前,对其进行全面评估至关重要,以确保其性能可靠。尽管学术界已充分认识到评估LLMs的重要性,但评估过程的复杂性导致了多样化的评估设置,进而引发研究结果与解读的不一致。为此,我们系统性地审视了在LLM评估各环节中导致这些不一致及不可靠评估的主要挑战与局限性。基于批判性分析,我们提出了确保LLM评估具备可复现性、可靠性与鲁棒性的观点与建议。