Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use. We study how TMK-based question generation strategies affect dataset quality for procedural and multi-hop reasoning. We compare three strategies: strict generation from Task-Method-Knowledge (TMK) models, transcript-first generation with post-hoc TMK filtering, and TMK-aware generation that combines transcripts with structured guidance. To evaluate generated items, we introduce a grounding validation framework based on closed-set evidence units extracted from TMK models. The framework measures whether answers are supported by the underlying representation, whether questions are self-contained, and whether they target multi-hop procedural reasoning. Across 23 instructional topics and 690 generated question-answer pairs, strict TMK generation achieves the strongest overall quality, with 96.5% grounded questions and 92.6% usable questions. Transcript-first generation produces more learner-like questions but more context-dependent or weakly grounded items, while TMK-aware generation yields high raw multi-hop coverage but lower grounding. These results show that procedural richness and natural phrasing do not guarantee representational grounding, motivating explicit representation-aware validation for evaluation datasets in AI-supported learning.
翻译:评估人工智能辅助学习系统中的程序性推理能力,需要既能模拟学习者特征、又能基于系统预期使用的教学知识的问答数据集。我们研究了基于TMK的问题生成策略对程序性推理和多跳推理数据集质量的影响。我们比较了三种策略:基于任务-方法-知识(TMK)模型的严格生成、基于转录优先生成结合事后TMK过滤、以及结合转录与结构化指导的TMK感知生成。为评估生成的条目,我们引入了一种基于TMK模型中提取的封闭式证据单元的基础性验证框架。该框架衡量答案是否得到底层表示的支持、问题是否自包含、以及是否针对多跳程序性推理。在23个教学主题和690对生成的问答对中,严格TMK生成获得了最强的整体质量,其中96.5%的问题具有基础性,92.6%的问题具有可用性。转录优先生成产生了更接近学习者的提问,但更多依赖于上下文或基础性较弱的条目,而TMK感知生成虽获得较高的原始多跳覆盖率,但基础性较低。这些结果显示,程序性丰富性和自然措辞并不能保证表示的基础性,从而推动了在人工智能辅助学习中的评估数据集需采用显式的表示感知验证方法。