Captions are crucial for understanding scientific visualizations and documents. Existing captioning methods for scientific figures rely on figure-caption pairs extracted from documents for training, many of which fall short with respect to metrics like helpfulness, explainability, and visual-descriptiveness [15] leading to generated captions being misaligned with reader preferences. To enable the generation of high-quality figure captions, we introduce FigCaps-HF a new framework for figure-caption generation that can incorporate domain expert feedback in generating captions optimized for reader preferences. Our framework comprises of 1) an automatic method for evaluating quality of figure-caption pairs, 2) a novel reinforcement learning with human feedback (RLHF) method to optimize a generative figure-to-caption model for reader preferences. We demonstrate the effectiveness of our simple learning framework by improving performance over standard fine-tuning across different types of models. In particular, when using BLIP as the base model, our RLHF framework achieves a mean gain of 35.7%, 16.9%, and 9% in ROUGE, BLEU, and Meteor, respectively. Finally, we release a large-scale benchmark dataset with human feedback on figure-caption pairs to enable further evaluation and development of RLHF techniques for this problem.
翻译:摘要: 标题对于理解科学可视化图表与文献至关重要。现有的科学图表标题生成方法依赖从文献中提取的图表-标题对进行训练,其中许多方法在有用性、可解释性和视觉描述性[15]等指标上表现不足,导致生成的标题与读者偏好不一致。为生成高质量的图表标题,本文提出FigCaps-HF这一新型图表标题生成框架,该框架能够整合领域专家反馈,生成针对读者偏好优化的标题。本框架包括:1) 一种自动评估图表-标题对质量的方法,2) 一种新颖的基于人类反馈的强化学习(RLHF)方法,以优化生成式图表到标题模型,使其符合读者偏好。我们通过在不同类型模型上提升性能优于标准微调,证明了该简单学习框架的有效性。具体而言,当使用BLIP作为基础模型时,本RLHF框架在ROUGE、BLEU和Meteor指标上分别获得了平均35.7%、16.9%和9%的提升。最后,我们发布了一个大规模基准数据集,包含图表-标题对的人类反馈,以支持该问题下RLHF技术的进一步评估与开发。