Within the millions of digitized historic American newspapers in the Chronicling America initiative are tens of millions of photographs, illustrations, cartoons, and advertisements. Much of this visual culture is shared across newspaper titles and issues. Just as reprinted texts within these newspapers speak to the virality of textual content, so too does this reprinted visual culture speak to newspapers as sites of constant information circulation and exchange. In this paper, we introduce Viral Images, a project to identify reprintings within 1.5 million photographs in Chronicling America. For our analysis, we adopt the Newspaper Navigator dataset of extracted photographs from over 16 million pages in Chronicling America. We introduce an unsupervised method of identifying reprintings by leveraging contrastive language-image pretraining (CLIP) to embed these 1.5 million photographs and applying clustering to identify re-printed content. We detail our public interface, https://viral-images.org, which we designed in order to enable humanists to interactively browse and study these identified clusters. In addition, we analyze the identified clusters, uncovering a diversity of photographs and advertisements that have been circulated across different newspapers over time.
翻译:在《纪事美国》倡议数字化的数百万份历史美国报纸中,包含了数千万张照片、插图、漫画和广告。这些视觉文化中有很大一部分跨越不同报纸标题和期号被共享。正如这些报纸中再版文本反映了文字内容的病毒式传播,这种再版视觉文化也表明报纸是信息不断流通与交换的场所。本文介绍了一项名为“病毒式图像”的研究项目,旨在从《纪事美国》的150万张照片中识别再版内容。为进行分析,我们采用了《报纸导航员》数据集,其中包含从《纪事美国》超过1600万页报纸中提取的照片。我们提出了一种无监督的再版识别方法,利用对比语言-图像预训练(CLIP)对150万张照片进行嵌入,并通过聚类分析识别再版内容。我们详细介绍了公共界面https://viral-images.org,该界面旨在让人文学者能够交互式浏览和研究这些识别的聚类。此外,我们对已识别的聚类进行了分析,揭示了在不同报纸上随时间传播的多样化照片和广告。