As the volume of data on the web grows, the web structure graph, which is a graph representation of the web, continues to evolve. The structure of this graph has gradually shifted from content-based to non-content-based. Furthermore, spam data, such as noisy hyperlinks, in the web structure graph adversely affect the speed and efficiency of information retrieval and link mining algorithms. Previous works in this area have focused on removing noisy hyperlinks using structural and string approaches. However, these approaches may incorrectly remove useful links or be unable to detect noisy hyperlinks in certain circumstances. In this paper, a data collection of hyperlinks is initially constructed using an interactive crawler. The semantic and relatedness structure of the hyperlinks is then studied through semantic web approaches and tools such as the DBpedia ontology. Finally, the removal process of noisy hyperlinks is carried out using a reasoner on the DBpedia ontology. Our experiments demonstrate the accuracy and ability of semantic web technologies to remove noisy hyperlinks
翻译:随着网络数据量的增长,作为网络图表示的网页结构图持续演化。该图的结构已逐渐从基于内容转向非内容导向。此外,网页结构图中的垃圾数据(如噪音超链接)会对信息检索与链接挖掘算法的速度及效率产生不利影响。现有研究主要采用结构化和字符串方法移除噪音超链接,但这些方法在某些情况下可能错误删除有效链接或无法识别噪音超链接。本文首先通过交互式爬虫构建超链接数据集合,随后利用语义网方法与工具(如DBpedia本体)研究超链接的语义与关联性结构,最后基于DBpedia本体使用推理器执行噪音超链接移除流程。实验结果表明,语义网技术能够准确有效地移除噪音超链接。