Dance serves as both a cultural cornerstone and a medium for personal expression, yet the rapid growth of online dance content has made personalized discovery increasingly difficult. Text-based dance retrieval offers a natural interface for users to search with choreographic intent, but it remains underexplored because dance requires simultaneous reasoning over linguistic semantics, musical rhythm, and full-body motion dynamics. We introduce TD-Data, a large-scale open dataset for text-dance retrieval, containing about 4,000 12-second dance clips, 14.6 hours of motion, 22 genres, and annotations from professional dance experts. On top of this dataset, we propose CustomDancer, a multimodal retrieval framework that aligns text with dance through a CLIP-based text encoder, music and motion encoders, and a music-motion blending module. CustomDancer achieves state-of-the-art performance on TD-Data, reaching 10.23% Recall@1 and improving retrieval quality in both quantitative benchmarks and user preference studies.
翻译:舞蹈既是文化的基石,也是个人表达的媒介,然而在线舞蹈内容的快速增长使得个性化发现日益困难。基于文本的舞蹈检索为用户提供了通过编舞意图进行搜索的自然交互界面,但由于舞蹈需要同时对语言语义、音乐节奏和全身运动动力学进行推理,该领域尚未得到充分探索。我们提出了TD-Data——一个面向文本-舞蹈检索的大规模开放数据集,包含约4000段12秒舞蹈片段、14.6小时运动数据、22种舞蹈流派及来自专业舞蹈专家的标注。在此数据集基础上,我们提出CustomDancer——一个多模态检索框架,通过基于CLIP的文本编码器、音乐与运动编码器以及音乐-运动融合模块实现文本与舞蹈的对齐。在TD-Data上,CustomDancer达到10.23%的Recall@1,在定量基准测试和用户偏好研究中均提升了检索质量。