We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of multimodal vision-language information of STEM. Our dataset features one of the largest and most comprehensive datasets for the challenge. It includes 448 skills and 1,073,146 questions spanning all STEM subjects. Compared to existing datasets that often focus on examining expert-level ability, our dataset includes fundamental skills and questions designed based on the K-12 curriculum. We also add state-of-the-art foundation models such as CLIP and GPT-3.5-Turbo to our benchmark. Results show that the recent model advances only help master a very limited number of lower grade-level skills (2.5% in the third grade) in our dataset. In fact, these models are still well below (averaging 54.7%) the performance of elementary students, not to mention near expert-level performance. To understand and increase the performance on our dataset, we teach the models on a training split of our dataset. Even though we observe improved performance, the model performance remains relatively low compared to average elementary students. To solve STEM problems, we will need novel algorithmic innovations from the community.
翻译:我们提出了一项新的挑战,旨在测试神经模型在STEM(科学、技术、工程与数学)领域的技能。现实世界中的问题往往需要结合STEM知识来寻求解决方案。与现有数据集不同,我们的数据集要求理解STEM领域的多模态视觉-语言信息。本数据集是该领域规模最大、最全面的数据集之一,涵盖448项技能和1,073,146个问题,覆盖所有STEM学科。相较于现有数据集通常侧重于考察专家级能力,我们的数据集基于K-12课程设计,包含基础技能与问题。我们还将CLIP和GPT-3.5-Turbo等先进基础模型纳入基准测试。结果显示,当前模型的最新进展仅能掌握我们数据集中极少数低年级技能(三年级技能的2.5%)。事实上,这些模型的性能(平均54.7%)仍远低于小学生水平,更遑论达到专家级表现。为理解和提升模型在此数据集上的性能,我们使用数据集的训练集对模型进行训练。尽管我们观察到性能有所提升,但模型表现仍相对低于普通小学生。要解决STEM问题,社区需要提出创新的算法方案。