Impressive advances in acquisition and sharing technologies have made the growth of multimedia collections and their applications almost unlimited. However, the opposite is true for the availability of labeled data, which is needed for supervised training, since such data is often expensive and time-consuming to obtain. While there is a pressing need for the development of effective retrieval and classification methods, the difficulties faced by supervised approaches highlight the relevance of methods capable of operating with few or no labeled data. In this work, we propose a novel manifold learning algorithm named Rank Flow Embedding (RFE) for unsupervised and semi-supervised scenarios. The proposed method is based on ideas recently exploited by manifold learning approaches, which include hypergraphs, Cartesian products, and connected components. The algorithm computes context-sensitive embeddings, which are refined following a rank-based processing flow, while complementary contextual information is incorporated. The generated embeddings can be exploited for more effective unsupervised retrieval or semi-supervised classification based on Graph Convolutional Networks. Experimental results were conducted on 10 different collections. Various features were considered, including the ones obtained with recent Convolutional Neural Networks (CNN) and Vision Transformer (ViT) models. High effective results demonstrate the effectiveness of the proposed method on different tasks: unsupervised image retrieval, semi-supervised classification, and person Re-ID. The results demonstrate that RFE is competitive or superior to the state-of-the-art in diverse evaluated scenarios.
翻译:采集与共享技术的显著进步使得多媒体集合及其应用的增长几乎不受限制。然而,标记数据的可用性却相反,这类数据对于监督训练至关重要,但获取代价高昂且耗时。尽管高效检索与分类方法的开发迫在眉睫,监督方法面临的困难凸显了能在少量或无标记数据下运行的方法的重要性。本文提出一种名为秩流嵌入的新型流形学习算法,适用于无监督与半监督场景。该方法基于近期流形学习研究中采用的概念,包括超图、笛卡尔积与连通分量。算法通过基于排名的处理流程逐步精炼上下文敏感嵌入,同时融入互补的上下文信息。生成的嵌入可用于更有效的无监督检索或基于图卷积网络的半监督分类。实验在10个不同数据集上进行,涵盖了包括基于卷积神经网络与视觉Transformer模型提取的多种特征。高效结果证明了该方法在不同任务上的有效性:无监督图像检索、半监督分类与行人重识别。结果表明RFE在多种评估场景中达到或超越现有最优方法。