One reason the Web is more useful than a simple collection of documents is that the structure created by hyperlinks enables flexible navigation from one web page to another. However, hyperlinks are typically created manually and cannot fully capture a corpus' implicit semantic structures. Is there a general way to make an arbitrary collection navigable? Recent work has formalized this problem generally as constructing a Hypergraph of Text (HoT), which provides a formal mathematical structure for supporting navigation and browsing. However, how to construct and evaluate a Hypergraph of Text remains a challenge. In this paper, we propose and study several methods for constructing a HoT. We also propose a novel quantitative metric, effort ratio, for evaluating the structural quality of a constructed HoT. Experimental results show that even simple TF-IDF baselines can match LLM-based methods on our proposed effort ratio metric.
翻译:万维网比简单文档集合更有用的一个原因在于,超链接创建的结构支持网页间的灵活导航。然而,超链接通常由人工创建,无法完整捕获语料库的隐含语义结构。是否存在一种通用方法能使任意集合具备可导航性?近期研究将该问题形式化为文本超图(Hypergraph of Text, HoT)的构建,该结构为支持导航与浏览提供了形式化数学框架。然而,如何构建和评估文本超图仍具挑战。本文提出并研究了若干种HoT构建方法,同时提出一种新的定量指标——努力比(effort ratio),用于评估所构建HoT的结构质量。实验结果表明,即使是简单的TF-IDF基线方法,也能在我们提出的努力比指标上达到与基于大语言模型的方法相当的效果。