Automated image captioning has the potential to be a useful tool for people with vision impairments. Images taken by this user group are often noisy, which leads to incorrect and even unsafe model predictions. In this paper, we propose a quality-agnostic framework to improve the performance and robustness of image captioning models for visually impaired people. We address this problem from three angles: data, model, and evaluation. First, we show how data augmentation techniques for generating synthetic noise can address data sparsity in this domain. Second, we enhance the robustness of the model by expanding a state-of-the-art model to a dual network architecture, using the augmented data and leveraging different consistency losses. Our results demonstrate increased performance, e.g. an absolute improvement of 2.15 on CIDEr, compared to state-of-the-art image captioning networks, as well as increased robustness to noise with up to 3 points improvement on CIDEr in more noisy settings. Finally, we evaluate the prediction reliability using confidence calibration on images with different difficulty/noise levels, showing that our models perform more reliably in safety-critical situations. The improved model is part of an assisted living application, which we develop in partnership with the Royal National Institute of Blind People.
翻译:自动图像描述有潜力成为视障人士的有用工具。该用户群体拍摄的图像往往存在噪声,这会导致模型产生错误甚至不安全的预测。本文提出一个质量无关框架,旨在提升面向视障人士的图像描述模型的性能与鲁棒性。我们从数据、模型和评估三个角度解决该问题。首先,我们展示如何通过生成合成噪声的数据增强技术来应对该领域的数据稀疏性。其次,我们将现有最优模型扩展为双网络架构,利用增强数据并引入多种一致性损失函数,从而增强模型鲁棒性。结果表明,与最优图像描述网络相比,我们的模型性能显著提升,例如CIDEr指标绝对提高2.15分,同时在噪声强度更高的场景下,CIDEr指标最多提升3分,展现出更强的噪声鲁棒性。最后,我们针对不同难度/噪声水平的图像,通过置信度校准评估预测可靠性,证明模型在安全关键场景中表现更可靠。改进后的模型已集成至与英国皇家盲人协会合作开发的辅助生活应用中。