Contrastive language-audio pretraining (CLAP) enables zero-shot audio classification, but standard inference classifies each clip in isolation and ignores the structure of the unlabeled test set. We present the first systematic study of TransCLIP-style transductive inference for CLAP: a text-anchored spherical Gaussian-mixture EM that refines zero-shot posteriors using the audio-embedding statistics of the test batch, with no labels, no gradients, and negligible compute (about 15 ms on one CPU core for 2,000 clips). Across ESC-50, UrbanSound8K, and VocalSound, this consistently improves top-1 accuracy by +4.6 to +9.2 points over the zero-shot baseline (e.g., 89.1 -> 94.8% on ESC-50, 73.8 -> 81.8% on UrbanSound8K). We further show that the gain (i) is governed by a simple operating boundary -- roughly 2.5 test samples per class per batch are required, with diminishing returns beyond ~5; (ii) is complementary to entropy-guided prompt weighting, with the combination reaching 96.2% on ESC-50; and (iii) attenuates but remains positive under long-tailed batches (+4.9 -> +3.1 points at a 20:1 imbalance), which we report as an explicit limitation. We also document a negative result: on TUT Urban Acoustic Scenes 2018, where zero-shot CLAP is near chance, transduction has no signal to amplify.
翻译:对比语言-音频预训练(CLAP)实现了零样本音频分类,但标准推理过程独立处理每个片段,忽略了未标注测试集的结构。我们首次系统研究了面向CLAP的TransCLIP风格直推式推理:一种以文本锚定的球面高斯混合期望最大化算法,该算法利用测试批次的音频嵌入统计信息优化零样本后验概率,无需标签、无需梯度,且计算量微乎其微(对2000个片段仅需单个CPU核心约15毫秒)。在ESC-50、UrbanSound8K和VocalSound数据集上,该方法相较于零样本基线在top-1准确率上持续提升4.6至9.2个百分点(例如,ESC-50从89.1%提升至94.8%,UrbanSound8K从73.8%提升至81.8%)。我们进一步证明:(i)该增益受简单操作边界控制——每个批次每类约需2.5个测试样本,超过约5个后收益递减;(ii)与熵引导提示加权互补,组合方法在ESC-50上达到96.2%;(iii)在长尾批次中增益减弱但仍保持正值(20:1不平衡度下从+4.9降至+3.1个百分点),我们将其作为明确局限性报告。我们还记录了一个负面结果:在TUT Urban Acoustic Scenes 2018数据集上,当零样本CLAP接近随机水平时,直推式方法无法增强有效信号。