Semantic dimensions of sound have been playing a central role in understanding the nature of auditory sensory experience as well as the broader relation between perception, language, and meaning. Accordingly, and given the recent proliferation of large language models (LLMs), here we asked whether such models exhibit an organisation of perceptual semantics similar to those observed in humans. Specifically, we prompted ChatGPT, a chatbot based on a state-of-the-art LLM, to rate musical instrument sounds on a set of 20 semantic scales. We elicited multiple responses in separate chats, analogous to having multiple human raters. ChatGPT generated semantic profiles that only partially correlated with human ratings, yet showed robust agreement along well-known psychophysical dimensions of musical sounds such as brightness (bright-dark) and pitch height (deep-high). Exploratory factor analysis suggested the same dimensionality but different spatial configuration of a latent factor space between the chatbot and human ratings. Unexpectedly, the chatbot showed degrees of internal variability that were comparable in magnitude to that of human ratings. Our work highlights the potential of LLMs to capture salient dimensions of human sensory experience.
翻译:声音的语义维度在理解听觉感官体验的本质以及感知、语言与意义之间的广泛关系中一直扮演着核心角色。因此,鉴于大型语言模型的近期普及,我们在此探究此类模型是否展现出与人类相似的感知语义组织。具体而言,我们引导ChatGPT——一个基于最先进大型语言模型的聊天机器人——在20个语义量表上对乐器声音进行评分。我们在独立的对话中收集了多个响应,类似于让多名人类评分者参与。ChatGPT生成的语义轮廓与人类评分仅部分相关,但在众所周知的音乐声音心理物理维度上显示出稳健的一致性,例如亮度(明亮-暗淡)和音高(低沉-高昂)。探索性因子分析表明,聊天机器人与人类评分之间具有相同的潜在维度,但潜在因子空间的空间配置不同。出乎意料的是,聊天机器人展现出的内部变异性程度与人类评分相当。我们的工作凸显了大型语言模型在捕捉人类感官体验显著维度方面的潜力。