Upon its release in late 2022, ChatGPT has brought a seismic shift in the entire landscape of AI, both in research and commerce. Through instruction-tuning a large language model (LLM) with supervised fine-tuning and reinforcement learning from human feedback, it showed that a model could answer human questions and follow instructions on a broad panel of tasks. Following this success, interests in LLMs have intensified, with new LLMs flourishing at frequent interval across academia and industry, including many start-ups focused on LLMs. While closed-source LLMs (e.g., OpenAI's GPT, Anthropic's Claude) generally outperform their open-source counterparts, the progress on the latter has been rapid with claims of achieving parity or even better on certain tasks. This has crucial implications not only on research but also on business. In this work, on the first anniversary of ChatGPT, we provide an exhaustive overview of this success, surveying all tasks where an open-source LLM has claimed to be on par or better than ChatGPT.
翻译:自2022年底发布以来,ChatGPT彻底改变了人工智能领域的整体格局,无论是研究还是商业领域均受其深刻影响。通过监督微调与基于人类反馈的强化学习对大型语言模型进行指令微调,该模型展现出在广泛任务中回答人类问题并遵循指令的能力。受此成功推动,大型语言模型领域的研究热度持续攀升,学术界与工业界以高频次涌现出大量新型LLM,其中不乏专注于该领域的初创企业。尽管闭源LLM(如OpenAI的GPT、Anthropic的Claude)通常优于开源同类模型,但后者进展迅猛,已在特定任务上声称达到甚至超越前者水平。这一态势对研究与商业均具有关键影响。值此ChatGPT发布一周年之际,本研究系统梳理了开源大语言模型宣称与ChatGPT持平或更优的所有任务,全面呈现这一里程碑式的成功。