Many information access systems operationalize their results in terms of rankings, which are then displayed to users in various ranking layouts such as linear lists or grids. User interaction with a retrieved item is highly dependent on the item's position in the layout, and users do not provide similar attention to every position in ranking (under any layout model). User attention is an important component in the evaluation process of ranking, due to its use in effectiveness metrics that estimate utility as well as fairness metrics that evaluate ranking based on social and ethical concerns. These metrics take user browsing behavior into account in their measurement strategies to estimate the attention the user is likely to provide to each item in ranking. Research on understanding user browsing behavior has proposed several user browsing models, and further observed that user browsing behavior differs with different ranking layouts. However, the underlying concepts of these browsing models are often similar, including varying components and parameter settings. We seek to leverage that similarity to represent multiple browsing models in a generalized, configurable framework which can be further extended to more complex ranking scenarios. In this paper, we describe a probabilistic user browsing model for linear rankings, show how they can be configured to yield models commonly used in current evaluation practice, and generalize this model to also account for browsing behaviors in grid-based layouts. This model provides configurable framework for estimating the attention that results from user browsing activity for a range of IR evaluation and measurement applications in multiple formats, and also identifies parameters that need to be estimated through user studies to provide realistic evaluation beyond ranked lists.
翻译:许多信息访问系统通过排序结果来运作,随后以线性列表或网格等不同排序布局向用户展示。用户对检索项目的交互高度依赖于项目在布局中的位置,且在任何布局模型下,用户对排序中每个位置的关注度并不相同。用户注意力是排序评估过程的重要组成部分,因为它被用于衡量效用的有效性指标以及基于社会与伦理考量评估排序的公平性指标。这些指标在其测量策略中纳入用户浏览行为,以估计用户可能对排序中每个项目给予的关注度。针对用户浏览行为的研究提出了多种用户浏览模型,并进一步观察到用户浏览行为因排序布局不同而存在差异。然而,这些浏览模型的基本概念往往相似,仅包含不同的组件和参数设置。我们旨在利用这种相似性,将多种浏览模型表示为一个通用、可配置的框架,该框架可进一步扩展至更复杂的排序场景。本文描述了一种用于线性排序的概率性用户浏览模型,展示了如何配置该模型以生成当前评估实践中常用的模型,并将该模型推广至基于网格的布局中的浏览行为。该模型提供了一个可配置框架,用于在多种格式的信息检索评估与测量应用中估计用户浏览活动产生的注意力,同时识别出需要通过用户研究来估计的参数,从而实现对排序列表之外场景的真实评估。