Advances in natural language processing, such as transfer learning from pre-trained language models, have impacted how models are trained for programming language tasks too. Previous research primarily explored code pre-training and expanded it through multi-modality and multi-tasking, yet the data for downstream tasks remain modest in size. Focusing on data utilization for downstream tasks, we propose and adapt augmentation methods that yield consistent improvements in code translation and summarization by up to 6.9% and 7.5% respectively. Further analysis suggests that our methods work orthogonally and show benefits in output code style and numeric consistency. We also discuss test data imperfections.
翻译:自然语言处理的进展(例如基于预训练语言模型的迁移学习)也影响了编程语言任务的模型训练方式。以往研究主要探索代码预训练,并通过多模态与多任务学习扩展其能力,然而下游任务的数据规模仍然有限。针对下游任务的数据利用问题,我们提出并适配了多种数据增强方法,在代码翻译与代码摘要任务中分别实现了高达6.9%和7.5%的持续改进。进一步分析表明,我们的方法具有正交性,并在输出代码风格与数值一致性方面展现出优势。此外,我们还讨论了测试数据中的不完善之处。