Do multilingual embedding models encode a language-general representation of proficiency? We investigate this by training linear and non-linear probes on hidden-state activations from Qwen3-Embedding (0.6B, 4B, 8B) to predict CEFR proficiency levels from learner texts across nine corpora and seven languages. We compare five probing architectures against a baseline trained on surface-level text features. Under in-distribution evaluation, probes achieve strong performance ($QWK\approx0.7$), substantially outperforming the surface baseline, with middle layers consistently yielding the best predictions. However, in cross-corpus evaluation performance collapses across all probe types and model sizes. Residual analysis reveals that out-of-distribution probes converge towards predicting uniformly distributed labels, indicating that the learned mappings capture corpus-specific distributional properties (topic, language, task type, rating methodology) rather than an abstract, transferable proficiency dimension. These results suggest that current multilingual embeddings do not straightforwardly encode language-general proficiency, with implications for representation-based approaches to proficiency-adaptive language technology.
翻译:多语言嵌入模型是否编码了语言通用的熟练度表示?我们通过从Qwen3-Embedding(0.6B、4B、8B)的隐状态激活中训练线性和非线性探针,来探究这一问题,并在跨越九个语料库和七种语言的CEFR熟练度等级预测任务上进行评估。我们比较了五种探针架构,并使用一个基于表层文本特征的基线模型作为对照。在分布内评估中,探针取得了强劲性能($QWK\approx0.7$),显著优于表层基线,其中中间层始终产生最佳预测。然而,在跨语料库评估中,所有探针类型和模型规模的性能均出现崩溃。残差分析表明,分布外探针趋向于预测均匀分布的标签,这表示所学到的映射捕获了语料库特定的分布特性(主题、语言、任务类型、评分方法),而非抽象的、可迁移的熟练度维度。这些结果表明,当前的多语言嵌入并未直接编码语言通用的熟练度,这对基于表示的熟练度自适应语言技术具有重要启示。