The Rasch model is one of the most fundamental models in \emph{item response theory} and has wide-ranging applications from education testing to recommendation systems. In a universe with $n$ users and $m$ items, the Rasch model assumes that the binary response $X_{li} \in \{0,1\}$ of a user $l$ with parameter $\theta^*_l$ to an item $i$ with parameter $\beta^*_i$ (e.g., a user likes a movie, a student correctly solves a problem) is distributed as $\Pr(X_{li}=1) = 1/(1 + \exp{-(\theta^*_l - \beta^*_i)})$. In this paper, we propose a \emph{new item estimation} algorithm for this celebrated model (i.e., to estimate $\beta^*$). The core of our algorithm is the computation of the stationary distribution of a Markov chain defined on an item-item graph. We complement our algorithmic contributions with finite-sample error guarantees, the first of their kind in the literature, showing that our algorithm is consistent and enjoys favorable optimality properties. We discuss practical modifications to accelerate and robustify the algorithm that practitioners can adopt. Experiments on synthetic and real-life datasets, ranging from small education testing datasets to large recommendation systems datasets show that our algorithm is scalable, accurate, and competitive with the most commonly used methods in the literature.
翻译:Rasch模型是项目反应理论中最基础的模型之一,广泛应用于教育测试到推荐系统等领域。在一个包含n个用户和m个项目的场景中,Rasch模型假设参数为θ*_l的用户l对参数为β*_i的项目i(例如,用户喜欢某部电影,学生正确解答某道题目)的二元响应X_{li} ∈ {0,1}服从分布Pr(X_{li}=1) = 1/(1 + exp{-(θ*_l - β*_i)})。本文针对这一经典模型提出了一种新的项目估计算法(即估计β*)。该算法的核心是计算定义在项目-项目图上的马尔可夫链的平稳分布。我们为算法贡献了有限样本误差保证,这是文献中首次提出的此类保证,表明算法具有一致性和良好的最优性性质。我们讨论了实践者可采用的实用改进措施,以加速和增强算法的鲁棒性。在从教育测试数据集到大型推荐系统数据集等合成数据和真实数据上的实验表明,我们的算法具有可扩展性、准确性,且与文献中最常用的方法相比具有竞争力。