Cybersecurity information is often technically complex and relayed through unstructured text, making automation of cyber threat intelligence highly challenging. For such text domains that involve high levels of expertise, pretraining on in-domain corpora has been a popular method for language models to obtain domain expertise. However, cybersecurity texts often contain non-linguistic elements (such as URLs and hash values) that could be unsuitable with the established pretraining methodologies. Previous work in other domains have removed or filtered such text as noise, but the effectiveness of these methods have not been investigated, especially in the cybersecurity domain. We propose different pretraining methodologies and evaluate their effectiveness through downstream tasks and probing tasks. Our proposed strategy (selective MLM and jointly training NLE token classification) outperforms the commonly taken approach of replacing non-linguistic elements (NLEs). We use our domain-customized methodology to train CyBERTuned, a cybersecurity domain language model that outperforms other cybersecurity PLMs on most tasks.
翻译:网络安全信息通常技术复杂且以非结构化文本形式传递,这使得网络威胁情报的自动化极具挑战性。对于这类涉及高专业知识的文本领域,领域内语料库预训练已成为语言模型获取领域专长的常用方法。然而,网络安全文本常包含URL、哈希值等非语言元素,这些元素可能与现有预训练方法不兼容。此前其他领域的研究中,这类元素常被作为噪声移除或过滤,但此类方法的有效性尚未得到充分验证,尤其在网络安全领域。我们提出了不同的预训练方法,并通过下游任务与探测任务评估其效果。我们所提出的策略(选择性掩码语言模型与联合训练非语言元素标记分类)优于常见的替换非语言元素方法。基于领域定制化方法,我们训练了CyBERTuned——一种在多数任务上优于其他网络安全预训练语言模型的网络安全领域语言模型。