In contemporary times, people rely heavily on the internet and search engines to obtain information, either directly or indirectly. However, the information accessible to users constitutes merely 4% of the overall information present on the internet, which is commonly known as the surface web. The remaining information that eludes search engines is called the deep web. The deep web encompasses deliberately hidden information, such as personal email accounts, social media accounts, online banking accounts, and other confidential data. The deep web contains several critical applications, including databases of universities, banks, and civil records, which are off-limits and illegal to access. The dark web is a subset of the deep web that provides an ideal platform for criminals and smugglers to engage in illicit activities, such as drug trafficking, weapon smuggling, selling stolen bank cards, and money laundering. In this article, we propose a search engine that employs deep learning to detect the titles of activities on the dark web. We focus on five categories of activities, including drug trading, weapon trading, selling stolen bank cards, selling fake IDs, and selling illegal currencies. Our aim is to extract relevant images from websites with a ".onion" extension and identify the titles of websites without images by extracting keywords from the text of the pages. Furthermore, we introduce a dataset of images called Darkoob, which we have gathered and used to evaluate our proposed method. Our experimental results demonstrate that the proposed method achieves an accuracy rate of 94% on the test dataset.
翻译:在当代,人们高度依赖互联网和搜索引擎直接或间接地获取信息。然而,用户可访问的信息仅占互联网信息总量的4%,这部分通常被称为表层网络。搜索引擎无法触及的其余信息被称为深网。深网包含刻意隐藏的信息,例如个人电子邮件账户、社交媒体账户、网上银行账户及其他机密数据。深网包含若干关键应用,如大学数据库、银行数据库和民事记录,这些信息禁止访问且访问属于违法行为。暗网是深网的一个子集,为犯罪分子和走私者提供了从事非法活动的理想平台,例如毒品交易、武器走私、出售被盗银行卡及洗钱。本文提出了一种搜索引擎,利用深度学习检测暗网上的活动标题。我们聚焦于五类活动:毒品交易、武器交易、出售被盗银行卡、出售假身份证及出售非法货币。我们的目标是从以“.onion”为扩展名的网站中提取相关图像,并通过提取页面文本中的关键词识别无图像网站的标题。此外,我们引入了一个名为Darkoob的图像数据集,该数据集由我们收集并用于评估所提出的方法。实验结果表明,所提方法在测试数据集上达到了94%的准确率。