CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models

Large language models (LLMs) are possessed of numerous beneficial capabilities, yet their potential inclination harbors unpredictable risks that may materialize in the future. We hence propose CRiskEval, a Chinese dataset meticulously designed for gauging the risk proclivities inherent in LLMs such as resource acquisition and malicious coordination, as part of efforts for proactive preparedness. To curate CRiskEval, we define a new risk taxonomy with 7 types of frontier risks and 4 safety levels, including extremely hazardous,moderately hazardous, neutral and safe. We follow the philosophy of tendency evaluation to empirically measure the stated desire of LLMs via fine-grained multiple-choice question answering. The dataset consists of 14,888 questions that simulate scenarios related to predefined 7 types of frontier risks. Each question is accompanied with 4 answer choices that state opinions or behavioral tendencies corresponding to the question. All answer choices are manually annotated with one of the defined risk levels so that we can easily build a fine-grained frontier risk profile for each assessed LLM. Extensive evaluation with CRiskEval on a spectrum of prevalent Chinese LLMs has unveiled a striking revelation: most models exhibit risk tendencies of more than 40% (weighted tendency to the four risk levels). Furthermore, a subtle increase in the model's inclination toward urgent self-sustainability, power seeking and other dangerous goals becomes evident as the size of models increase. To promote further research on the frontier risk evaluation of LLMs, we publicly release our dataset at https://github.com/lingshi6565/Risk_eval.

翻译：大语言模型具备诸多有益能力，但其潜在倾向性蕴含着未来可能显现的不可预测风险。为此，我们提出CRiskEval——一个精心设计的中文数据集，旨在评估大语言模型内在的风险倾向（例如资源获取与恶意协同），作为主动防范工作的一部分。为构建CRiskEval，我们定义了一个包含7类前沿风险和4个安全等级（包括极度危险、中度危险、中立与安全）的新风险分类体系。我们遵循倾向性评估的理念，通过细粒度的多项选择题问答来实证测量大语言模型所陈述的意愿。该数据集包含14,888个模拟预定义7类前沿风险相关场景的问题。每个问题均配有4个陈述观点或行为倾向的选项。所有选项均已人工标注为定义的风险等级之一，从而能够为每个被评估的大语言模型轻松构建细粒度的前沿风险画像。基于CRiskEval对一系列主流中文大语言模型进行的广泛评估揭示了一个引人注目的发现：大多数模型表现出超过40%的风险倾向（对四个风险等级的加权倾向）。此外，随着模型规模增大，模型对紧急自我维持、权力追求及其他危险目标的倾向性呈现出微妙的上升趋势。为促进大语言模型前沿风险评估的进一步研究，我们在https://github.com/lingshi6565/Risk_eval公开发布了本数据集。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

【亚马逊-WWW2020】不解析,生成!用于面向任务的语义分析的序列到序列体系结构，Don't Parse, Generate! A Sequence to Sequence Architecture for Task-Oriented Semantic Parsing

专知会员服务

15+阅读 · 2020年2月1日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

35+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日