Deep code generation is a topic of deep learning for software engineering (DL4SE), which adopts neural models to generate code for the intended functions. Since end-to-end neural methods lack domain knowledge and software hierarchy awareness, they tend to perform poorly w.r.t project-level tasks. To systematically explore the potential improvements of code generation, we let it participate in the whole top-down development from \emph{expressibles} to \emph{executables}, which is possible in limited scopes. In the process, it benefits from massive samples, features, and knowledge. As the foundation, we suggest building a taxonomy on code data, namely code taxonomy, leveraging the categorization of code information. Moreover, we introduce a three-layer semantic pyramid (SP) to associate text data and code data. It identifies the information of different abstraction levels, and thus introduces the domain knowledge on development and reveals the hierarchy of software. Furthermore, we propose a semantic pyramid framework (SPF) as the approach, focusing on software of high modularity and low complexity. SPF divides the code generation process into stages and reserves spots for potential interactions. In addition, we conceived preliminary applications in software development to confirm the neuro-symbolic framework.
翻译:深度代码生成是软件工程深度学习领域(DL4SE)的一个研究方向,它采用神经模型为预期功能生成代码。由于端到端神经方法缺乏领域知识和软件层次意识,它们在项目级任务上表现较差。为系统探索代码生成的潜在改进空间,我们将其纳入从*可表述*到*可执行*的完整自上而下开发流程中,这在有限范围内是可行的。在此过程中,它得益于海量样本、特征和知识。作为基础,我们建议构建代码数据的分类体系,即代码分类法,利用代码信息的分类。此外,我们引入三层语义金字塔(SP)来关联文本数据和代码数据。该金字塔识别不同抽象层次的信息,从而引入开发领域知识并揭示软件层次结构。进一步,我们提出语义金字塔框架(SPF)作为方法,重点关注高模块化、低复杂度的软件。SPF将代码生成过程划分为多个阶段,并为潜在交互预留位置。此外,我们构思了软件开发中的初步应用以验证该神经符号框架。