Textbooks are the primary vehicle for delivering quality education to students. It has been shown that explanatory or illustrative visuals play a key role in the retention, comprehension and the general transfer of knowledge. However, many textbooks, especially in the developing world, are low quality and lack interesting visuals to support student learning. In this paper, we investigate the effectiveness of vision-language models to automatically enhance textbooks with images from the web. Specifically, we collect a dataset of e-textbooks from one of the largest free online publishers in the world. We rigorously analyse the dataset, and use the resulting analysis to motivate a task that involves retrieving and appropriately assigning web images to textbooks, which we frame as a novel optimization problem. Through a crowd-sourced evaluation, we verify that (1) while the original textbook images are rated higher, automatically assigned ones are not far behind, and (2) the choice of the optimization problem matters. We release the dataset of textbooks with an associated image bank to spur further research in this area.
翻译:教材是向学生传递优质教育的主要载体。已有研究表明,解释性或说明性图像在知识保持、理解及一般性知识迁移中发挥着关键作用。然而,许多教材(尤其是发展中国家的教材)质量较低,缺乏能支持学生学习的趣味性图像。本文探讨了利用视觉-语言模型自动从网络获取图像以增强教材的有效性。具体而言,我们从全球最大的免费在线出版商之一收集了电子教材数据集,对该数据集进行了严格分析,并基于分析结果提出了一项任务:涉及从网络检索并合理分配图像到教材中。我们将此任务定义为一个新型优化问题。通过众包评估,我们验证了:(1)尽管原始教材图像评分更高,但自动分配的图像评分差距不大;(2)优化问题的选择至关重要。我们公开了该教材数据集及关联图像库,以推动该领域的进一步研究。