Responding to the thousands of student questions on online QA platforms each semester has a considerable human cost, particularly in computing courses with rapidly growing enrollments. To address the challenges of scalable and intelligent question-answering (QA), we introduce an innovative solution that leverages open-source Large Language Models (LLMs) from the LLaMA-2 family to ensure data privacy. Our approach combines augmentation techniques such as retrieval augmented generation (RAG), supervised fine-tuning (SFT), and learning from human preferences data using Direct Preference Optimization (DPO). Through extensive experimentation on a Piazza dataset from an introductory CS course, comprising 10,000 QA pairs and 1,500 pairs of preference data, we demonstrate a significant 30% improvement in the quality of answers, with RAG being a particularly impactful addition. Our contributions include the development of a novel architecture for educational QA, extensive evaluations of LLM performance utilizing both human assessments and LLM-based metrics, and insights into the challenges and future directions of educational data processing. This work paves the way for the development of AI-TA, an intelligent QA assistant customizable for courses with an online QA platform
翻译:每学期在在线问答平台上回复学生数千个问题需要大量人力成本,尤其在注册人数快速增长的计算课程中这一问题尤为突出。为应对可扩展且智能的问答(QA)挑战,我们提出了一种创新解决方案,该方案利用来自LLaMA-2系列的开源大语言模型(LLMs)以确保数据隐私。我们的方法结合了检索增强生成(RAG)、监督微调(SFT)以及使用直接偏好优化(DPO)从人类偏好数据中学习等增强技术。通过在包含10,000个问答对和1,500个偏好数据对的计算机科学入门课程Piazza数据集上进行广泛实验,我们展示了答案质量显著提升30%,其中RAG作为特别有效的增强手段。我们的贡献包括:开发面向教育问答的新型架构、利用人工评估和基于大语言模型的指标对LLM性能进行广泛评估,以及对教育数据处理中的挑战与未来方向的深入洞察。这项工作为开发AI-TA(一种可针对在线问答平台课程定制的智能问答助手)奠定了基础。