The task of generating code from a natural language description, or NL2Code, is considered a pressing and significant challenge in code intelligence. Thanks to the rapid development of pre-training techniques, surging large language models are being proposed for code, sparking the advances in NL2Code. To facilitate further research and applications in this field, in this paper, we present a comprehensive survey of 27 existing large language models for NL2Code, and also review benchmarks and metrics. We provide an intuitive comparison of all existing models on the HumanEval benchmark. Through in-depth observation and analysis, we provide some insights and conclude that the key factors contributing to the success of large language models for NL2Code are "Large Size, Premium Data, Expert Tuning". In addition, we discuss challenges and opportunities regarding the gap between models and humans. We also create a website https://nl2code.github.io to track the latest progress through crowd-sourcing. To the best of our knowledge, this is the first survey of large language models for NL2Code, and we believe it will contribute to the ongoing development of the field.
翻译:从自然语言描述生成代码(即NL2Code)的任务被认为是代码智能领域一个紧迫且重要的挑战。得益于预训练技术的快速发展,针对代码的海量大型语言模型不断涌现,推动了NL2Code领域的进步。为促进该领域的进一步研究与应用,本文对27个现有NL2Code大型语言模型进行了全面综述,并回顾了相关基准测试与评价指标。我们在HumanEval基准上对所有现有模型进行了直观比较。通过深入观察与分析,我们提出了一些见解,并总结出NL2Code大型语言模型成功的关键因素为“大规模、优质数据、专家调优”。此外,我们探讨了模型与人类之间差距所带来的挑战与机遇。我们还创建了网站https://nl2code.github.io,通过众包方式追踪最新进展。据我们所知,这是首篇针对NL2Code大型语言模型的综述,我们相信它将为该领域的持续发展做出贡献。