Existing visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and precise concept of the target object as images are limited in their ability to abstract and can be ambiguous, which makes it difficult to track targets with drastic variations. In this paper, we propose the CiteTracker to enhance target modeling and inference in visual tracking by connecting images and text. Specifically, we develop a text generation module to convert the target image patch into a descriptive text containing its class and attribute information, providing a comprehensive reference point for the target. In addition, a dynamic description module is designed to adapt to target variations for more effective target representation. We then associate the target description and the search image using an attention-based correlation module to generate the correlated features for target state reference. Extensive experiments on five diverse datasets are conducted to evaluate the proposed algorithm and the favorable performance against the state-of-the-art methods demonstrates the effectiveness of the proposed tracking method.
翻译:现有视觉跟踪方法通常以目标图像块作为跟踪参照。然而,单一图像块因抽象能力有限且存在歧义,无法提供完整精准的目标对象概念,导致难以跟踪剧烈变化的目标。本文提出CiteTracker,通过连接图像与文本增强视觉跟踪中的目标建模与推理能力。具体而言,我们开发文本生成模块将目标图像块转换为包含类别与属性信息的描述性文本,为目标提供全面的参考基点。此外,设计动态描述模块适应目标变化,实现更有效的目标表征。进而采用基于注意力机制的关联模块,将目标描述与搜索图像进行关联以生成用于目标状态参考的关联特征。在五个不同数据集上的大量实验验证了所提算法的有效性,其相较于现有最优方法的优越性能证明了该跟踪方法的有效性。