Computerized Adaptive Testing (CAT) is a widely used, efficient test mode that adapts to the examinee's proficiency level in the test domain. CAT requires pre-trained item profiles, for CAT iteratively assesses the student real-time based on the registered items' profiles, and selects the next item to administer using candidate items' profiles. However, obtaining such item profiles is a costly process that involves gathering a large, dense item-response data, then training a diagnostic model on the collected data. In this paper, we explore the possibility of leveraging response data collected in the CAT service. We first show that this poses a unique challenge due to the inherent selection bias introduced by CAT, i.e., more proficient students will receive harder questions. Indeed, when naively training the diagnostic model using CAT response data, we observe that item profiles deviate significantly from the ground-truth. To tackle the selection bias issue, we propose the user-wise aggregate influence function method. Our intuition is to filter out users whose response data is heavily biased in an aggregate manner, as judged by how much perturbation the added data will introduce during parameter estimation. This way, we may enhance the performance of CAT while introducing minimal bias to the item profiles. We provide extensive experiments to demonstrate the superiority of our proposed method based on the three public datasets and one dataset that contains real-world CAT response data.
翻译:计算机化自适应测试(CAT)是一种广泛使用且高效的测试模式,能够根据考生在测试领域的能力水平进行自适应调整。CAT需要预训练的项目配置文件,因为CAT会基于已注册项目的配置文件实时评估学生的能力,并利用候选项目的配置文件选择下一个要施测的项目。然而,获取此类项目配置文件的过程成本高昂,涉及收集大规模、密集的项目-响应数据,并基于收集到的数据训练诊断模型。本文探索了利用CAT服务中收集的响应数据的可能性。我们首先表明,由于CAT固有的选择偏差(即能力较高的学生会收到更难的题目),这会带来独特的挑战。实际上,当使用CAT响应数据朴素地训练诊断模型时,我们观察到项目配置文件与真实情况存在显著偏差。为了解决选择偏差问题,我们提出了用户级聚合影响函数方法。我们的直觉是,以聚合方式过滤掉那些响应数据高度偏差的用户,判断依据是新增数据在参数估计过程中引入的扰动程度。通过这种方式,我们可以在增强CAT性能的同时,尽可能降低对项目配置文件的偏差引入。我们基于三个公开数据集和一个包含真实CAT响应数据的数据集进行了大量实验,证明了所提出方法的优越性。