We introduce C-Pack, a package of resources that significantly advance the field of general Chinese embeddings. C-Pack includes three critical resources. 1) C-MTEB is a comprehensive benchmark for Chinese text embeddings covering 6 tasks and 35 datasets. 2) C-MTP is a massive text embedding dataset curated from labeled and unlabeled Chinese corpora for training embedding models. 3) C-TEM is a family of embedding models covering multiple sizes. Our models outperform all prior Chinese text embeddings on C-MTEB by up to +10% upon the time of the release. We also integrate and optimize the entire suite of training methods for C-TEM. Along with our resources on general Chinese embedding, we release our data and models for English text embeddings. The English models achieve state-of-the-art performance on MTEB benchmark; meanwhile, our released English data is 2 times larger than the Chinese data. All these resources are made publicly available at https://github.com/FlagOpen/FlagEmbedding.
翻译:我们推出C-Pack,这是一套显著推进通用中文嵌入领域的资源包。C-Pack包含三大关键资源:1) C-MTEB——覆盖6类任务、35个数据集的中文文本嵌入综合基准;2) C-MTP——基于标注与非标注中文语料构建的大规模文本嵌入数据集,用于训练嵌入模型;3) C-TEM——覆盖多种规模的嵌入模型系列。在发布时,我们的模型在C-MTEB上的表现优于所有此前的中文文本嵌入模型,提升幅度最高达+10%。我们还将C-TEM的全套训练方法进行了整合与优化。伴随通用中文嵌入资源,我们还发布了英文文本嵌入的数据与模型。英文模型在MTEB基准上达到当前最优性能;同时,我们发布的英文数据规模是中文数据的两倍。以上所有资源均已公开于https://github.com/FlagOpen/FlagEmbedding。