As Large Language Models (LLMs) become increasingly integrated into everyday life, their capabilities to understand and emulate human cognition are under steady examination. This study investigates the ability of LLMs to comprehend and interpret linguistic pragmatics, an aspect of communication that considers context and implied meanings. Using Grice's communication principles, LLMs and human subjects (N=76) were evaluated based on their responses to various dialogue-based tasks. The findings revealed the superior performance and speed of LLMs, particularly GPT4, over human subjects in interpreting pragmatics. GPT4 also demonstrated accuracy in the pre-testing of human-written samples, indicating its potential in text analysis. In a comparative analysis of LLMs using human individual and average scores, the models exhibited significant chronological improvement. The models were ranked from lowest to highest score, with GPT2 positioned at 78th place, GPT3 ranking at 23rd, Bard at 10th, GPT3.5 placing 5th, Best Human scoring 2nd, and GPT4 achieving the top spot. The findings highlight the remarkable progress made in the development and performance of these LLMs. Future studies should consider diverse subjects, multiple languages, and other cognitive aspects to fully comprehend the capabilities of LLMs. This research holds significant implications for the development and application of AI-based models in communication-centered sectors.
翻译:随着大型语言模型(LLMs)日益融入日常生活,其理解和模拟人类认知的能力正受到持续检验。本研究调查了LLMs理解与阐释语言语用学(一种考虑语境和隐含意义的交流维度)的能力。基于格赖斯沟通原则,我们通过评估LLMs和人类受试者(N=76)对各种基于对话任务的反应展开研究。结果表明,LLMs,尤其是GPT-4,在语用阐释方面展现出优于人类受试者的性能和速度。GPT-4在人类书面样本的预测试中也表现出准确性,凸显了其在文本分析中的潜力。在基于人类个体和平均分数的LLMs比较分析中,模型呈现显著的时序性能提升。模型按得分从低到高排序依次为:GPT2位列第78名,GPT3排名第23名,Bard位列第10名,GPT3.5位居第5名,最佳人类得分者排名第2名,而GPT-4占据榜首。这些发现凸显了LLMs开发与性能上的显著进步。未来研究应考虑多样化受试者、多语种及其他认知维度,以全面理解LLMs的能力。本研究对以通信为核心的领域中的AI模型开发与应用具有重要启示意义。