We present a study of Tip-of-the-tongue (ToT) retrieval for music, where a searcher is trying to find an existing music entity, but is unable to succeed as they cannot accurately recall important identifying information. ToT information needs are characterized by complexity, verbosity, uncertainty, and possible false memories. We make four contributions. (1) We collect a dataset - $ToT_{Music}$ - of 2,278 information needs and ground truth answers. (2) We introduce a schema for these information needs and show that they often involve multiple modalities encompassing several Music IR subtasks such as lyric search, audio-based search, audio fingerprinting, and text search. (3) We underscore the difficulty of this task by benchmarking a standard text retrieval approach on this dataset. (4) We investigate the efficacy of query reformulations generated by a large language model (LLM), and show that they are not as effective as simply employing the entire information need as a query - leaving several open questions for future research.
翻译:我们针对音乐检索中的舌尖效应(Tip-of-the-Tongue, ToT)现象开展研究,在该场景中,搜索者试图查找某个已知的音乐实体,但因无法准确回忆起关键标识信息而检索失败。此类 ToT 信息需求具有复杂性、冗长性、不确定性以及可能的错误记忆等特征。我们做出四项贡献:(1)构建包含 2,278 条信息需求及其真实答案的数据集 $ToT_{Music}$;(2)针对这些信息需求提出分类体系,揭示其常涉及多模态特征,涵盖歌词搜索、基于音频的搜索、音频指纹识别及文本搜索等多项音乐信息检索(Music IR)子任务;(3)通过在该数据集上对标准文本检索方法进行基准测试,凸显该任务的难度;(4)探究由大语言模型(LLM)生成的查询改写方案的效果,结果表明其并不优于直接将整个信息需求用作查询——这为未来研究留下了若干待解问题。