We introduce CHARM, the first benchmark for comprehensively and in-depth evaluating the commonsense reasoning ability of large language models (LLMs) in Chinese, which covers both globally known and Chinese-specific commonsense. We evaluated 7 English and 12 Chinese-oriented LLMs on CHARM, employing 5 representative prompt strategies for improving LLMs' reasoning ability, such as Chain-of-Thought. Our findings indicate that the LLM's language orientation and the task's domain influence the effectiveness of the prompt strategy, which enriches previous research findings. We built closely-interconnected reasoning and memorization tasks, and found that some LLMs struggle with memorizing Chinese commonsense, affecting their reasoning ability, while others show differences in reasoning despite similar memorization performance. We also evaluated the LLMs' memorization-independent reasoning abilities and analyzed the typical errors. Our study precisely identified the LLMs' strengths and weaknesses, providing the clear direction for optimization. It can also serve as a reference for studies in other fields. We will release CHARM at https://github.com/opendatalab/CHARM .
翻译:我们提出了CHARM,首个系统深入评估大语言模型中文常识推理能力的基准测试,涵盖全球通用常识与中文特有常识。我们在CHARM上评估了7个英文和12个中文大语言模型,采用了5种具有代表性的提示策略(如思维链)以提升模型的推理能力。研究发现,模型的语言偏好与任务领域会影响提示策略的有效性,这一发现丰富了现有研究成果。我们构建了紧密关联的推理与记忆任务,发现部分模型在记忆中文常识方面存在困难从而影响其推理能力,而另一些模型尽管记忆表现相似,但推理能力存在差异。我们还评估了模型独立于记忆的推理能力,并分析了典型错误类型。本研究精准识别了模型的优势与不足,为优化方向提供了明确指引,同时可为其他领域研究提供参考。我们将于https://github.com/opendatalab/CHARM 公开发布CHARM。