Data similarity (or distance) computation is a fundamental research topic which fosters a variety of similarity-based machine learning and data mining applications. In big data analytics, it is impractical to compute the exact similarity of data instances due to high computational cost. To this end, the Locality Sensitive Hashing (LSH) technique has been proposed to provide accurate estimators for various similarity measures between sets or vectors in an efficient manner without the learning process. Structured data (e.g., sequences, trees and graphs), which are composed of elements and relations between the elements, are commonly seen in the real world, but the traditional LSH algorithms cannot preserve the structure information represented as relations between elements. In order to conquer the issue, researchers have been devoted to the family of the hierarchical LSH algorithms. In this paper, we explore the present progress of the research into hierarchical LSH from the following perspectives: 1) Data structures, where we review various hierarchical LSH algorithms for three typical data structures and uncover their inherent connections; 2) Applications, where we review the hierarchical LSH algorithms in multiple application scenarios; 3) Challenges, where we discuss some potential challenges as future directions.
翻译:数据相似性(或距离)计算是一个基础研究课题,支撑着多种基于相似性的机器学习与数据挖掘应用。在大数据分析中,由于计算代价高昂,无法计算数据实例的精确相似性。为此,局部敏感哈希(LSH)技术被提出,能够在无需学习过程的情况下,高效地为集合或向量间的多种相似性度量提供精确估计。结构化数据(如序列、树和图)由元素及元素间关系组成,在现实世界中普遍存在,但传统LSH算法无法保留以元素间关系表示的结构信息。为解决该问题,研究者致力于层级LSH算法家族的研究。本文从以下角度探讨层级LSH的研究进展:1) 数据结构方面,回顾针对三种典型数据结构的各类层级LSH算法,并揭示其内在联系;2) 应用方面,梳理层级LSH算法在多种应用场景中的运用;3) 挑战方面,讨论若干潜在挑战作为未来研究方向。