Alignment has become a critical step for instruction-tuned Large Language Models (LLMs) to become helpful assistants. However, effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluations tailored for alignment. To fill in this gap, we introduce AlignBench, a comprehensive multi-dimensional benchmark for evaluating LLMs' alignment in Chinese. Equipped with a human-in-the-loop data curation pipeline, our benchmark employs a rule-calibrated multi-dimensional LLM-as-Judge with Chain-of-Thought to generate explanations and final ratings as evaluations, ensuring high reliability and interpretability. Furthermore, we developed a dedicated companion evaluator LLM -- CritiqueLLM, which recovers 95\% of GPT-4's evaluation ability and will be provided via public APIs to researchers for evaluation of alignment in Chinese LLMs. All evaluation codes, data, and LLM generations are available at \url{https://github.com/THUDM/AlignBench}.
翻译:对齐已成为指令调优的大型语言模型(LLMs)成为有用助手的关键步骤。然而,针对新兴中文LLMs对齐效果的有效评估仍严重缺失,亟需基于真实场景的、开放式的、具有挑战性且自动化的对齐评估方法。为填补这一空白,我们提出了AlignBench——一个用于评估LLMs中文对齐能力的综合性多维基准。该基准采用人机协同的数据筛选流程,并利用基于规则校准的多维LLM-as-Judge方法,结合思维链(Chain-of-Thought)生成评估解释与最终评分,确保高可靠性与可解释性。此外,我们开发了专用辅助评估模型CritiqueLLM,该模型可恢复GPT-4 95%的评估能力,并将通过公共API向研究人员开放,用于中文LLMs对齐评估。所有评估代码、数据及LLM生成结果均可在\url{https://github.com/THUDM/AlignBench}获取。