Despite being proposed as early as 1959, COBOL (Common Business-Oriented Language) still predominantly acts as an integral part of the majority of operations of several financial, banking, and governmental organizations. To support the inevitable modernization and maintenance of legacy systems written in COBOL, it is essential for organizations, researchers, and developers to understand the nature and source code of COBOL programs. However, to the best of our knowledge, we are unaware of any dataset that provides data on COBOL software projects, motivating the need for the dataset. Thus, to aid empirical research on comprehending COBOL in open-source repositories, we constructed a dataset of 84 COBOL repositories mined from GitHub, containing rich metadata on the development cycle of the projects. We envision that researchers can utilize our dataset to study COBOL projects' evolution, code properties and develop tools to support their development. Our dataset also provides 1255 COBOL files present inside the mined repositories. The dataset and artifacts are available at https://doi.org/10.5281/zenodo.7968845.
翻译:尽管COBOL(面向商业的通用语言)早在1959年就被提出,但它目前仍是多家金融、银行和政府机构大多数业务运营的核心组成部分。为支持用COBOL编写的遗留系统不可避免的现代化改造与维护,组织、研究人员和开发者亟需理解COBOL程序的性质及源代码。然而,据我们所知,目前尚无提供COBOL软件项目相关数据的数据集,这凸显了构建该数据集的必要性。为此,我们构建了包含84个从GitHub挖掘的COBOL仓库的数据集,其中涵盖项目开发周期的丰富元数据,以支持对开源仓库中COBOL代码的实证研究。我们预期研究人员可借助该数据集研究COBOL项目的演进规律、代码特性,并开发相关工具以支持其开发。该数据集还包含挖掘所得仓库中的1255个COBOL文件。数据集及工件可通过https://doi.org/10.5281/zenodo.7968845获取。