When many people highlight the same document, is the crowd a single consensus, or is it internally structured into reader sub-groups that mark different things -- and is that structure a stable property of a reader or of the document? Building on prior work showing an individual's within-document highlighting signal is a whisper while individuality lives in selection, we ask the group-level question on a co-readership platform using a margin-preserving curveball null. Experiment 1: within a document, readers form strong sub-groups -- pairs agree far beyond what shared salience, mark density, and sentence popularity predict (nearest-neighbour agreement z=+6.3, significant in 88% of documents). Under an eight-block region-preserving null, shared engagement with the same coarse regions of the document accounts for about 40% of this excess; the majority survives as finer reader-specific agreement (z=+3.6, 77% significant). So the within-document crowd is, in a descriptive sense, factional. Experiment 2: is that grouping a stable reader trait? Here we are honest about power. The cross-document split-half reproducibility of a pair's agreement is near zero pooled (+0.078 and 0.000 in two separately drawn samples), and a power calibration shows the test is informative only for pairs that co-read many documents. In the only informative high-overlap subset (k>=4), point estimates are positive but small-sample, imprecise across the separately drawn samples, never significant, and attenuate under the region-preserving null. We therefore leave cross-document stability unresolved: the data is consistent with anything from situational grouping to a weak-to-moderate stable reader trait. The crowd is factional within a document; whether its factions follow the reader across documents is, honestly, beyond our reach.
翻译:当众多读者高亮同一文档时,这些读者是形成单一共识,还是内部结构化为标记不同内容的读者子群?这种结构是读者的稳定属性,还是文档的稳定属性?基于先前研究(个体在文档内的高亮信号如同细语,而个性体现在选择中),我们在一共读平台上提出群体层面的问题,并采用保留边界的curveball零假设。实验一:在文档内,读者形成强子群——成对读者的一致性远超共享显著度、标记密度和句子流行度所预测的水平(最近邻一致性z=+6.3,在88%的文档中显著)。在八区块保留区域的零假设下,对文档相同粗粒度区域的共同参与解释了约40%的额外一致性;剩余大部分表现为更精细的读者特定一致性(z=+3.6,77%显著)。因此,在描述意义上,文档内读者群呈现派系化。实验二:这种分组是否为稳定的读者特质?此处我们坦诚统计效力。跨文档的分半信度显示,成对读者一致性的可重复性近乎零(在两个独立抽取样本中分别为+0.078和0.000),而效力校准表明,该检验仅对共读多个文档的成对读者具有信息性。在唯一信息性的高重叠子集(k≥4)中,点估计为正但小样本,在独立抽取样本间不精确,从未显著,且在保留区域零假设下衰减。因此,我们暂不解决跨文档稳定性问题:数据与从情境分组到弱至中稳定读者特质的各种可能性均一致。读者群在文档内呈现派系化;但其派系是否随读者跨越文档——坦言之,超出我们能力范围。