Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingual ToT test collections for Chinese, Japanese, Korean, and English, using an LLM-based query simulation framework. We systematically study how prompt language and source document language affect the fidelity of simulated ToT queries, validating synthetic queries through system rank correlation against real user queries. Our results show that effective ToT simulation requires language-aware design choices: non-English language sources are generally important, while English Wikipedia can be beneficial when non-English sources provide insufficient information for query generation. Based on these findings, we release four ToT test collections with 5,000 queries per language across multiple domains. This work provides the first large-scale multilingual ToT benchmark and offers practical guidance for constructing realistic ToT datasets beyond English.
翻译:舌尖效应(Tip-of-the-Tongue, ToT)检索基准主要集中在英语,限制了其在多语言信息检索中的适用性。本文利用基于大语言模型的查询模拟框架,构建了中文、日语、韩语及英语的多语言ToT测试集。我们系统研究了提示语言与源文档语言对模拟ToT查询逼真度的影响,并通过系统排名与真实用户查询的相关性验证合成查询的有效性。结果表明,有效的ToT模拟需要语言感知的设计选择:非英语语言源通常至关重要,而当非英语源无法为查询生成提供足够信息时,英语维基百科则可发挥有益作用。基于这些发现,我们发布了四个ToT测试集,每种语言包含5000个跨多个领域的查询。本研究提供了首个大规模多语言ToT基准,并为构建英语以外的逼真ToT数据集提供了实用指导。