Model Construction is a foundational practice in science learning that relies on visualization and interactivity. Large Language Models, increasingly augmented with multimodal capabilities, have been integrated in education contexts to support learning. However, these tools lack visual interactivity that is required by some learning contexts. We introduce Text to Multimodal Model (T2MM), a robust, dynamic LLM supported architecture that assists in model construction within the open inquiry ecology-based modeling software Virtual Experimental Research Assistant (VERA). T2MM accounts for the current context of the learner's model and creates interactive models, rather than static images, enabling the model to remain responsive to manual adjustment. To measure technical feasibility, we evaluate T2MM through a custom procedurally generated dataset of natural language learner modeling requests and target models within the VERA system. T2MM outperforms a baseline model generation architecture implemented through LLM-supported full code generation, common in the literature, across all measured success metrics. Our contribution not only outlines LLM integration into a inquiry-based learning modeling tool, but also describes a possible architecture through which more interactive multimodal LLM tools can be created.
翻译:模型构建是科学学习中的基础实践,依赖于可视化与交互性。近年来,具有多模态能力增强的大型语言模型已被整合到教育场景中以支持学习。然而,这些工具缺乏某些学习场景所需的视觉交互性。我们提出文本到多模态模型(T2MM),这是一种稳健、动态的LLM支持架构,可在开放探究生态建模软件虚拟实验研究助手(VERA)中辅助模型构建。T2MM能够捕捉学习者模型当前上下文,创建交互式模型(而非静态图像),从而使模型保持对人工调整的响应性。为衡量技术可行性,我们通过VERA系统中自定义程序生成的自然语言学习者建模请求及目标模型数据集对T2MM进行评估。在所有评估指标上,T2MM均优于文献中常见的基于LLM全代码生成的基线模型生成架构。我们的贡献不仅概述了LLM在探究式学习建模工具中的集成方式,还描述了一种可创建更具交互性的多模态LLM工具的潜在架构。