The Competition on Software Verification (SV-COMP) is a large computational experiment benchmarking many different software verification tools on a vast collection of C and Java benchmarks. Such experimental research should be reproducible by researchers independent from the team that performed the original experiments. In this reproduction report, we present our recent attempt at reproducing SV-COMP 2023: We chose a meaningful subset of the competition and re-ran it on the competition organiser's infrastructure, using the scripts and tools provided in the competition's archived artifacts. We see minor differences in tool scores that appear explainable by the interaction of small runtime fluctuations with the competition's scoring rules, and successfully reproduce the overall ranking within our chosen subset. Overall, we consider SV-COMP 2023 to be reproducible.
翻译:软件验证竞赛(SV-COMP)是一项大规模计算实验,旨在对大量C和Java基准测试集上的多种软件验证工具进行基准测试。此类实验研究应能由独立于原始实验团队的研究人员复现。本复现报告中,我们介绍了近期对SV-COMP 2023的复现尝试:我们从竞赛中选取有意义的子集,在竞赛组织者的基础设施上使用竞赛归档制品中提供的脚本和工具重新运行实验。我们发现工具得分存在微小差异,这些差异可通过小幅度运行时波动与竞赛评分规则的相互作用得到合理解释,并成功复现了所选子集的整体排名。总体而言,我们认为SV-COMP 2023具有可复现性。