Rising cyber threats, with miscreants registering thousands of new domains daily for Internet-scale attacks like spam, phishing, and drive-by downloads, emphasize the need for innovative detection methods. This paper introduces a cutting-edge approach for identifying suspicious domains at the onset of the registration process. The accompanying data pipeline generates crucial features by comparing new domains to registered domains,emphasizing the crucial similarity score. Leveraging a novel combination of Natural Language Processing (NLP) techniques, including a pretrained Canine model, and Multilayer Perceptron (MLP) models, our system analyzes semantic and numerical attributes, providing a robust solution for early threat detection. This integrated approach significantly reduces the window of vulnerability, fortifying defenses against potential threats. The findings demonstrate the effectiveness of the integrated approach and contribute to the ongoing efforts in developing proactive strategies to mitigate the risks associated with illicit online activities through the early identification of suspicious domain registrations.
翻译:日益增长的网络安全威胁——不法分子每天注册数千个新域名用于垃圾邮件、网络钓鱼和驱动下载等大规模互联网攻击——凸显了创新检测方法的必要性。本文提出了一种在注册过程起始阶段识别可疑域名的前沿方法。配套的数据管道通过将新域名与已注册域名进行比较来生成关键特征,其中重点强调了关键相似度评分。通过利用自然语言处理(NLP)技术(包括预训练的Canine模型)与多层感知器(MLP)模型的新颖组合,我们的系统能够分析语义和数值属性,为早期威胁检测提供鲁棒解决方案。这种集成方法显著缩短了脆弱性窗口期,强化了针对潜在威胁的防御能力。研究结果证明了该集成方法的有效性,并通过早期识别可疑域名注册,为制定主动策略以减轻非法网络活动相关风险的持续努力做出了贡献。