Although much of the success of Deep Learning builds on learning good representations, a rigorous method to evaluate their quality is lacking. In this paper, we treat the evaluation of representations as a model selection problem and propose to use the Minimum Description Length (MDL) principle to devise an evaluation metric. Contrary to the established practice of limiting the capacity of the readout model, we design a hybrid discrete and continuous-valued model space for the readout models and employ a switching strategy to combine their predictions. The MDL score takes model complexity, as well as data efficiency into account. As a result, the most appropriate model for the specific task and representation will be chosen, making it a unified measure for comparison. The proposed metric can be efficiently computed with an online method and we present results for pre-trained vision encoders of various architectures (ResNet and ViT) and objective functions (supervised and self-supervised) on a range of downstream tasks. We compare our methods with accuracy-based approaches and show that the latter are inconsistent when multiple readout models are used. Finally, we discuss important properties revealed by our evaluations such as model scaling, preferred readout model, and data efficiency.
翻译:尽管深度学习的大部分成功建立在学习良好表示的基础上,但缺乏评估其质量的严谨方法。本文将表示评估视为模型选择问题,并提出利用最小描述长度原则设计评估指标。与传统限制读出模型容量的做法不同,我们为读出模型设计了一个混合离散与连续值的模型空间,并采用切换策略来组合其预测结果。MDL评分同时考虑了模型复杂度和数据效率,因此能够为特定任务和表示选择最合适的模型,从而成为统一的比较度量。所提出的指标可通过在线方法高效计算,我们展示了针对不同架构(ResNet与ViT)和优化目标(监督与自监督)的预训练视觉编码器在多个下游任务上的结果。我们将所提方法与基于准确率的评估方法进行比较,发现后者在使用多个读出模型时存在不一致性。最后,我们讨论了评估所揭示的重要特性,例如模型规模扩展、偏好读出模型以及数据效率。