This paper presents a modular AI agentic skill pipeline for automating subject indexing with Library of Congress Subject Headings (LCSH). Subject indexing - the process of analyzing a work's aboutness, selecting controlled vocabulary terms, and encoding them as MARC21 subject access fields - is one of the most time-consuming components of library cataloging. The system decomposes this process into four discrete, sequentially executed agent skills: conceptual analysis, quantitative filtering, authority validation, and MARC field synthesis. Each skill encodes domain knowledge drawn directly from Library of Congress Subject Headings Manual (SHM) instruction sheets and subject analysis theory. The pipeline was evaluated against a corpus of ten titles whose existing subject headings were captured from the Harvard Library bibliographic dataset (a snapshot of their Alma ILS). Results demonstrate strong conceptual alignment with professional subject indexing practice, with notable differences in specificity, subdivision practice, and the agent's adherence to the 2026 LC policy discontinuing form subdivisions in favor of LCGFT 655 fields.
翻译:本文提出了一种模块化的AI主体技能流水线,用于自动化执行美国国会图书馆主题标引(Library of Congress Subject Headings, LCSH)。主题标引——即分析作品的主题内容、选择受控词汇术语并将其编码为MARC21主题检索字段的过程——是图书馆编目中最耗时的环节之一。该系统将该流程分解为四个独立的、按序执行的智能体技能:概念分析、定量过滤、权威验证和MARC字段合成。每项技能直接编码来自《美国国会图书馆主题标引手册》(SHM)指导说明及主题分析理论中的领域知识。该流水线基于十种文献的语料库进行评估,这些文献的现有主题标引从哈佛图书馆书目数据集(其Alma ILS的快照)中获取。结果表明,该流水线与专业主题标引实践在概念上高度一致,但在具体性、副标引实践及智能体遵循2026年美国国会图书馆政策(即用LCGFT 655字段取代形式副标引)方面存在显著差异。