We present BookReconciler, an open-source tool for enhancing and clustering book data. BookReconciler allows users to take spreadsheets with minimal metadata, such as book title and author, and automatically 1) add authoritative, persistent identifiers like ISBNs 2) and cluster related Expressions and Manifestations of the same Work, e.g., different translations or editions. This enhancement makes it easier to combine related collections and analyze books at scale. The tool is currently designed as an extension for OpenRefine -- a popular software application -- and connects to major bibliographic services including the Library of Congress, VIAF, OCLC, HathiTrust, Google Books, and Wikidata. Our approach prioritizes human judgment. Through an interactive interface, users can manually evaluate matches and define the contours of a Work (e.g., to include translations or not). We evaluate reconciliation performance on datasets of U.S. prize-winning books and contemporary world fiction. BookReconciler achieves near-perfect accuracy for U.S. works but lower performance for global texts, reflecting structural weaknesses in bibliographic infrastructures for non-English and global literature. Overall, BookReconciler supports the reuse of bibliographic data across domains and applications, contributing to ongoing work in digital libraries and digital humanities.
翻译:我们提出BookReconciler,一款用于增强与聚类图书数据的开源工具。该工具支持用户导入仅含书名、作者等基础元数据的电子表格,自动完成两项核心功能:1)添加权威持久标识符(如ISBN);2)对同一作品的关联表现形式与具体版本(如不同译本或版本)进行聚类。这种增强处理便于合并相关馆藏并进行规模化图书分析。当前工具被设计为流行软件OpenRefine的扩展插件,可对接美国国会图书馆、VIAF、OCLC、HathiTrust、Google Books及Wikidata等主要书目服务系统。我们的方法强调人工判断:通过交互界面,用户能够手动评估匹配结果,定义作品边界(如是否纳入译本)。我们以美国获奖图书及当代世界虚构类作品数据集评估校对性能。结果显示,BookReconciler对美国作品实现近乎完美的准确率,但对全球性文本表现欠佳——这暴露了非英语与全球文学在书目基础设施领域的结构性缺陷。总体而言,BookReconciler促进了跨领域与跨应用场景的书目数据复用,为数字图书馆与数字人文领域的持续研究作出贡献。