School choice mechanism designers use discrete choice models to understand and predict families' preferences. The most widely-used choice model, the multinomial logit (MNL), is linear in school and/or household attributes. While the model is simple and interpretable, it assumes the ranked preference lists arise from a choice process that is uniform throughout the ranking, from top to bottom. In this work, we introduce two strategies for rank-heterogeneous choice modeling tailored for school choice. First, we adapt a context-dependent random utility model (CDM), considering down-rank choices as occurring in the context of earlier up-rank choices. Second, we consider stratifying the choice modeling by rank, regularizing rank-adjacent models towards one another when appropriate. Using data on household preferences from the San Francisco Unified School District (SFUSD) across multiple years, we show that the contextual models considerably improve our out-of-sample evaluation metrics across all rank positions over the non-contextual models in the literature. Meanwhile, stratifying the model by rank can yield more accurate first-choice predictions while down-rank predictions are relatively unimproved. These models provide performance upgrades that school choice researchers can adopt to improve predictions and counterfactual analyses.
翻译:学校选择机制设计者使用离散选择模型来理解和预测家庭的偏好。应用最广泛的选择模型——多项式逻辑模型(MNL),在学区属性和/或家庭属性上呈线性关系。尽管该模型简单且可解释,但它假设排名偏好列表源于从高到低整个排名过程中一致的选择过程。在本研究中,我们提出了两种针对学校选择的秩异质性选择建模策略。首先,我们引入了一个情境依赖随机效用模型(CDM),将低排名选择视为在高排名选择情境下发生。其次,我们考虑按排名分层建模,并在适当时对相邻秩模型进行正则化。利用来自旧金山联合学区(SFUSD)多年间的家庭偏好数据,我们证明情境模型在所有排名位置上均显著优于文献中的非情境模型,提高了样本外评估指标。同时,按排名分层模型可提供更准确的第一选择预测,而低排名预测则改善有限。这些模型为学校选择研究者提供了性能提升方案,可用于改进预测及反事实分析。