Changes in satellite imagery often occur over multiple time steps. Despite the emergence of bi-temporal change captioning datasets, there is a lack of multi-temporal event captioning datasets (at least two images per sequence) in remote sensing. This gap exists because (1) searching for visible events in satellite imagery and (2) labeling multi-temporal sequences require significant time and labor. To address these challenges, we present SkyScraper, an iterative multi-agent workflow that geocodes news articles and synthesizes captions for corresponding satellite image sequences. Our experiments show that SkyScraper successfully finds 5x more events than traditional geocoding methods, demonstrating that agentic feedback is an effective strategy for surfacing new multi-temporal events in satellite imagery. We apply our framework to a large database of global news articles, curating a new multi-temporal captioning dataset with 5,000 sequences. By automatically identifying imagery related to news events, our work also supports journalism and reporting efforts.
翻译:卫星图像中的变化通常涉及多个时间步长。尽管双时相变化描述数据集已出现,但遥感领域仍缺乏多时相事件描述数据集(每个序列至少包含两幅图像)。这一空白存在的原因在于:(1)在卫星图像中搜索可见事件;(2)标注多时相序列需要大量时间与人力。为解决上述挑战,我们提出SkyScraper——一种迭代式多智能体工作流,可对新闻文章进行地理编码并合成对应卫星图像序列的描述文本。实验表明,与传统地理编码方法相比,SkyScraper成功发现的事件数量提升5倍,证明智能体反馈策略在卫星图像中挖掘新型多时相事件方面具有显著效果。我们将该框架应用于大规模全球新闻文章数据库,整理出包含5000个序列的新型多时相描述数据集。通过自动识别与新闻事件相关的图像,本研究同时为新闻业与报道工作提供了技术支撑。