Remote sensing visual question answering (RSVQA) opens new opportunities for the use of overhead imagery by the general public, by enabling human-machine interaction with natural language. Building on the recent advances in natural language processing and computer vision, the goal of RSVQA is to answer a question formulated in natural language about a remote sensing image. Language understanding is essential to the success of the task, but has not yet been thoroughly examined in RSVQA. In particular, the problem of language biases is often overlooked in the remote sensing community, which can impact model robustness and lead to wrong conclusions about the performances of the model. Thus, the present work aims at highlighting the problem of language biases in RSVQA with a threefold analysis strategy: visual blind models, adversarial testing and dataset analysis. This analysis focuses both on model and data. Moreover, we motivate the use of more informative and complementary evaluation metrics sensitive to the issue. The gravity of language biases in RSVQA is then exposed for all of these methods with the training of models discarding the image data and the manipulation of the visual input during inference. Finally, a detailed analysis of question-answer distribution demonstrates the root of the problem in the data itself. Thanks to this analytical study, we observed that biases in remote sensing are more severe than in standard VQA, likely due to the specifics of existing remote sensing datasets for the task, e.g. geographical similarities and sparsity, as well as a simpler vocabulary and question generation strategies. While new, improved and less-biased datasets appear as a necessity for the development of the promising field of RSVQA, we demonstrate that more informed, relative evaluation metrics remain much needed to transparently communicate results of future RSVQA methods.
翻译:遥感视觉问答(RSVQA)通过自然语言实现人机交互,为公众利用遥感图像开辟了新机遇。该任务借助自然语言处理与计算机视觉的最新进展,旨在回答关于遥感图像的自然语言问题。语言理解对该任务的成功至关重要,但在RSVQA中尚未得到充分研究。尤其值得一提的是,遥感领域常忽视语言偏向问题,这会影响模型鲁棒性并导致对模型性能的错误判断。因此,本文通过三种分析策略(视觉盲模型、对抗性测试和数据集分析)聚焦RSVQA中的语言偏向问题,从模型与数据两个层面展开研究。此外,我们倡导采用对这一问题更敏感、更具信息量且互补的评估指标。通过训练忽略图像数据的模型以及推理时操纵视觉输入,我们揭示了RSVQA中语言偏向的严重性。最后,对问答分布的详细分析表明问题根源在于数据本身。通过这一分析研究,我们发现遥感中的偏向比标准VQA更为严重,这很可能源于现有遥感数据集在该任务上的特殊性,例如地理相似性与稀疏性,以及更简化的词汇与问题生成策略。尽管开发更新、更优且偏向性更低的数据集是推动RSVQA这一前景领域发展的必要条件,但我们的研究表明,未来RSVQA方法仍需更具信息量的相对评估指标,以透明地传达其结果。