Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and instructional processes. However, traditional qualitative analysis remains a labor-intensive bottleneck, severely limiting the scale at which this research can be conducted. We present Sandpiper, a mixed-initiative system designed to serve as a bridge between high-volume conversational data and human qualitative expertise. By tightly coupling interactive researcher dashboards with agentic Large Language Model (LLM) engines, the platform enables scalable analysis without sacrificing methodological rigor. Sandpiper addresses critical barriers to AI adoption in education by implementing context-aware, automated de-identification workflows supported by secure, university-housed infrastructure to ensure data privacy. Furthermore, the system employs schema-constrained orchestration to eliminate LLM hallucinations and enforces strict adherence to qualitative codebooks. An integrated evaluations engine allows for the continuous benchmarking of AI performance against human labels, fostering an iterative approach to model refinement and validation. We propose a user study to evaluate the system's efficacy in improving research efficiency, inter-rater reliability, and researcher trust in AI-assisted qualitative workflows.
翻译:摘要:数字教育环境正朝着复杂的人工智能与人类话语融合的方向发展,为研究者提供了蕴含学习与教学过程深层洞察的海量数据。然而,传统定性分析仍是劳动密集型瓶颈,严重制约了此类研究的规模化开展。我们提出Sandpiper——一种混合主动式系统,旨在成为高容量对话数据与人类定性专长之间的桥梁。通过将交互式研究者仪表盘与智能体大型语言模型(LLM)引擎紧密耦合,该平台在保证方法论严谨性的前提下实现可扩展分析。Sandpiper通过实施情境感知的自动化去标识化工作流程,配合大学托管的安全基础设施确保数据隐私,从而解决教育领域AI应用的关键障碍。此外,系统采用模式约束编排消除LLM幻觉,强制严格遵循定性编码手册。集成评估引擎可对AI标注与人工标签进行持续基准测试,推动模型优化与验证的迭代方法。我们拟开展用户研究评估该系统在提升研究效率、评分者间信度及研究者对AI辅助定性工作流信任度方面的效能。