The widespread accessibility of the Internet has led to a surge in online fraudulent activities, underscoring the necessity of shielding users' sensitive information from cybercriminals. Phishing, a well-known cyberattack, revolves around the creation of phishing webpages and the dissemination of corresponding URLs, aiming to deceive users into sharing their sensitive information, often for identity theft or financial gain. Various techniques are available for preemptively categorizing zero-day phishing URLs by distilling unique attributes and constructing predictive models. However, these existing techniques encounter unresolved issues. This proposal delves into persistent challenges within phishing detection solutions, particularly concentrated on the preliminary phase of assembling comprehensive datasets, and proposes a potential solution in the form of a tool engineered to alleviate bias in ML models. Such a tool can generate phishing webpages for any given set of legitimate URLs, infusing randomly selected content and visual-based phishing features. Furthermore, we contend that the tool holds the potential to assess the efficacy of existing phishing detection solutions, especially those trained on confined datasets.
翻译:互联网的广泛普及导致网络欺诈活动激增,保护用户敏感信息免受网络犯罪分子侵害的必要性日益凸显。网络钓鱼作为一种典型网络攻击,主要通过创建钓鱼网页并传播对应URL,诱骗用户泄露身份信息或造成财产损失。现有技术可通过提取独特属性并构建预测模型,对零日钓鱼URL进行预分类。然而,这些技术仍存在未解决的关键问题。本研究深入探讨了钓鱼检测解决方案中持续存在的挑战,尤其聚焦于构建综合数据集的初始阶段,并提出了一种潜在解决方案——设计用于缓解机器学习模型偏差的工具。该工具可为任意给定合法URL生成钓鱼网页,通过随机注入内容与基于视觉的钓鱼特征。此外,我们认为该工具还有望评估现有钓鱼检测方案(尤其是基于封闭数据集训练的模型)的有效性。