Visual active search (VAS) has been proposed as a modeling framework in which visual cues are used to guide exploration, with the goal of identifying regions of interest in a large geospatial area. Its potential applications include identifying hot spots of rare wildlife poaching activity, search-and-rescue scenarios, identifying illegal trafficking of weapons, drugs, or people, and many others. State of the art approaches to VAS include applications of deep reinforcement learning (DRL), which yield end-to-end search policies, and traditional active search, which combines predictions with custom algorithmic approaches. While the DRL framework has been shown to greatly outperform traditional active search in such domains, its end-to-end nature does not make full use of supervised information attained either during training, or during actual search, a significant limitation if search tasks differ significantly from those in the training distribution. We propose an approach that combines the strength of both DRL and conventional active search by decomposing the search policy into a prediction module, which produces a geospatial distribution of regions of interest based on task embedding and search history, and a search module, which takes the predictions and search history as input and outputs the search distribution. We develop a novel meta-learning approach for jointly learning the resulting combined policy that can make effective use of supervised information obtained both at training and decision time. Our extensive experiments demonstrate that the proposed representation and meta-learning frameworks significantly outperform state of the art in visual active search on several problem domains.
翻译:视觉主动搜索(VAS)被提出作为一种建模框架,利用视觉线索引导搜索,旨在识别大规模地理区域中的感兴趣区域。其潜在应用包括识别稀有野生动物偷猎活动的热点、搜救场景、识别武器、毒品或人口的非法贩运等。目前VAS的前沿方法包括深度强化学习(DRL)的应用(可生成端到端搜索策略)以及传统主动搜索(将预测与定制算法方法相结合)。尽管DRL框架在此类领域已被证明大幅优于传统主动搜索,但其端到端特性未能充分利用训练或实际搜索过程中获取的监督信息——当搜索任务与训练分布差异显著时,这一局限尤为突出。我们提出一种结合DRL与常规主动搜索优势的方法,通过将搜索策略分解为预测模块(基于任务嵌入和搜索历史生成感兴趣区域的地理空间分布)和搜索模块(以预测结果和搜索历史为输入,输出搜索分布)。我们开发了一种新颖的元学习方法,用于联合学习整合后的策略,使其能够有效利用训练和决策时获取的监督信息。大量实验表明,所提出的表征与元学习框架在多个问题域中显著优于视觉主动搜索的前沿方法。