Bridge and Hint: Extending Pre-trained Language Models for Long-Range Code

In the field of code intelligence, effectively modeling long-range code poses a significant challenge. Existing pre-trained language models (PLMs) such as UniXcoder have achieved remarkable success, but they still face difficulties with long code inputs. This is mainly due to their limited capacity to maintain contextual continuity and memorize the key information over long-range code. To alleviate the difficulties, we propose EXPO, a framework for EXtending Pre-trained language models for lOng-range code. EXPO incorporates two innovative memory mechanisms we propose in this paper: Bridge Memory and Hint Memory. Bridge Memory uses a tagging mechanism to connect disparate snippets of long-range code, helping the model maintain contextual coherence. Hint Memory focuses on crucial code elements throughout the global context, such as package imports, by integrating a kNN attention layer to adaptively select the relevant code elements. This dual-memory approach bridges the gap between understanding local code snippets and maintaining global code coherence, thereby enhancing the model overall comprehension of long code sequences. We validate the effectiveness of EXPO on five popular pre-trained language models such as UniXcoder and two code intelligence tasks including API recommendation and vulnerability detection. Experimental results demonstrate that EXPO significantly improves the pre-training language models.

翻译：在代码智能领域，有效建模长程代码是一项重大挑战。现有预训练语言模型（如UniXcoder）虽已取得显著成功，但在处理长代码输入时仍面临困难，主要原因是其维持上下文连贯性及在长程代码中记忆关键信息的能力有限。为缓解这些问题，我们提出EXPO框架——一种扩展预训练语言模型以处理长程代码的框架。EXPO集成了本文提出的两种创新性记忆机制：桥接记忆与提示记忆。桥接记忆通过标签机制连接长程代码中分散的片段，帮助模型保持上下文连贯性；提示记忆则聚焦全局上下文中的关键代码元素（如包导入），通过集成k近邻注意力层自适应选择相关代码元素。这种双记忆方法填补了理解局部代码片段与维持全局代码连贯性之间的鸿沟，从而增强模型对长代码序列的整体理解能力。我们在五种主流预训练语言模型（如UniXcoder）及两项代码智能任务（包括API推荐与漏洞检测）上验证了EXPO的有效性。实验结果表明，EXPO显著提升了预训练语言模型的性能。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日