Language models have been increasingly popular for automatic creativity assessment, generating semantic distances to objectively measure the quality of creative ideas. However, there is currently a lack of an automatic assessment system for evaluating creative ideas in the Chinese language. To address this gap, we developed TransDis, a scoring system using transformer-based language models, capable of providing valid originality (quality) and flexibility (variety) scores for Alternative Uses Task (AUT) responses in Chinese. Study 1 demonstrated that the latent model-rated originality factor, comprised of three transformer-based models, strongly predicted human originality ratings, and the model-rated flexibility strongly correlated with human flexibility ratings as well. Criterion validity analyses indicated that model-rated originality and flexibility positively correlated to other creativity measures, demonstrating similar validity to human ratings. Study 2 & 3 showed that TransDis effectively distinguished participants instructed to provide creative vs. common uses (Study 2) and participants instructed to generate ideas in a flexible vs. persistent way (Study 3). Our findings suggest that TransDis can be a reliable and low-cost tool for measuring idea originality and flexibility in Chinese language, potentially paving the way for automatic creativity assessment in other languages. We offer an open platform to compute originality and flexibility for AUT responses in Chinese and over 50 other languages (https://osf.io/59jv2/).
翻译:语言模型在自动创造力评估中日益普及,通过生成语义距离来客观衡量创意想法的质量。然而,目前尚缺乏用于评估中文创意想法的自动评估系统。为填补这一空白,我们开发了TransDis——一个基于Transformer语言模型的评分系统,能够为中文替代用途测试(AUT)的回答提供有效的独创性(质量)和灵活性(多样性)评分。研究1表明,由三个基于Transformer模型构成的潜在模型评分独创性因子能有力预测人类独创性评分,且模型评分灵活性与人类灵活性评分高度相关。效标效度分析显示,模型评定的独创性和灵活性与其它创造力测量指标正相关,展现出与人类评分相当的效度。研究2和3表明,TransDis能有效区分被要求提供创造性用途与常见用途的参与者(研究2),以及被要求以灵活方式与持续方式生成想法的参与者(研究3)。我们的研究结果表明,TransDis可作为测量中文想法独创性和灵活性的可靠低成本工具,并可能为其他语言的自动创造力评估铺平道路。我们提供了一个开放平台,用于计算中文及其他50多种语言AUT回答的独创性和灵活性分数(https://osf.io/59jv2/)。