We present Spacerini, a modular framework for seamless building and deployment of interactive search applications, designed to facilitate the qualitative analysis of large scale research datasets. Spacerini integrates features from both the Pyserini toolkit and the Hugging Face ecosystem to ease the indexing text collections and deploy them as search engines for ad-hoc exploration and to make the retrieval of relevant data points quick and efficient. The user-friendly interface enables searching through massive datasets in a no-code fashion, making Spacerini broadly accessible to anyone looking to qualitatively audit their text collections. This is useful both to IR~researchers aiming to demonstrate the capabilities of their indexes in a simple and interactive way, and to NLP~researchers looking to better understand and audit the failure modes of large language models. The framework is open source and available on GitHub: https://github.com/castorini/hf-spacerini, and includes utilities to load, pre-process, index, and deploy local and web search applications. A portfolio of applications created with Spacerini for a multitude of use cases can be found by visiting https://hf.co/spacerini.
翻译:我们提出Spacerini,一个用于无缝构建和部署交互式搜索应用的模块化框架,旨在促进大规模研究数据集的定性分析。Spacerini整合了Pyserini工具包和Hugging Face生态系统的功能,以简化文本集合的索引过程,并将其部署为搜索引擎,用于临时探索,同时实现相关数据点的快速高效检索。其用户友好型界面支持以无代码方式搜索海量数据集,使得任何希望对其文本集合进行定性审查的用户均可广泛使用Spacerini。这一工具既有助于信息检索(IR)研究人员以简单互动的方式展示其索引能力,也有助于自然语言处理(NLP)研究人员更好地理解并审查大型语言模型的失效模式。该框架为开源项目,可在GitHub(https://github.com/castorini/hf-spacerini)上获取,并包含用于加载、预处理、索引及部署本地与网络搜索应用的实用工具。通过访问https://hf.co/spacerini,可查看使用Spacerini创建的多场景应用案例集。