Since the release of ChatGPT, numerous studies have highlighted the remarkable performance of ChatGPT, which often rivals or even surpasses human capabilities in various tasks and domains. However, this paper presents a contrasting perspective by demonstrating an instance where human performance excels in typical tasks suited for ChatGPT, specifically in the domain of computer programming. We utilize the IEEExtreme Challenge competition as a benchmark, a prestigious, annual international programming contest encompassing a wide range of problems with different complexities. To conduct a thorough evaluation, we selected and executed a diverse set of 102 challenges, drawn from five distinct IEEExtreme editions, using three major programming languages: Python, Java, and C++. Our empirical analysis provides evidence that contrary to popular belief, human programmers maintain a competitive edge over ChatGPT in certain aspects of problem-solving within the programming context. In fact, we found that the average score obtained by ChatGPT on the set of IEEExtreme programming problems is 3.9 to 5.8 times lower than the average human score, depending on the programming language. This paper elaborates on these findings, offering critical insights into the limitations and potential areas of improvement for AI-based language models like ChatGPT.
翻译:自ChatGPT发布以来,诸多研究凸显了其在各类任务与领域中媲美甚至超越人类能力的卓越表现。然而,本文通过展现一个与典型ChatGPT应用场景(即计算机编程领域)中人类表现更优的实例,提出了截然不同的观点。我们以享有盛誉的年度国际编程竞赛——IEEExtreme挑战赛为基准,该赛事涵盖复杂度各异的大量编程问题。为进行严谨评估,我们从五届不同的IEEExtreme赛事中精选并执行了102道代表性题目,涉及Python、Java和C++三种主流编程语言。实验分析表明,与普遍认知相反,人类程序员在编程问题求解的某些方面仍保持对ChatGPT的竞争优势。事实上,我们发现ChatGPT在IEEExtreme编程题集中的平均得分仅为人类平均得分的3.9至5.8分之一(具体因编程语言而异)。本文详细阐述了这些发现,为ChatGPT等基于人工智能的语言模型提供了关于其局限性及潜在改进方向的关键见解。